{
  "id": 354713,
  "title": "What are non \"CD**\" targets in CITE-seq? (like 'Mouse-IgG2a', 'Mouse-IgG2b',    'Rat-IgG2b' )",
  "url": "/competitions/open-problems-multimodal/discussion/354713",
  "author_name": "Alexander Chervov",
  "post_date": "2022-09-23T13:12:30.767000",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Would you be so kind to comment what these means : <br>\n'HLA-A-B-C', 'TIGIT', 'Mouse-IgG1', 'Mouse-IgG2a', 'Mouse-IgG2b',<br>\n   'Rat-IgG2b', 'Podoplanin', 'IgM', 'KLRG1', 'HLA-DR', 'CX3CR1',<br>\n   'integrinB7', 'TCR', 'Rat-IgG1', 'Rat-IgG2a', 'FceRIa', 'IgD',<br>\n   'TCRVa7.2', 'TCRVd2', 'LOX-1', 'HLA-E'</p>\n<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
  "messages": [
    {
      "id": 1952448,
      "postDate": "2022-09-23T17:14:02.977Z",
      "content": "<p>You can find full information about the antibody panel for the CITE here: <a href=\"https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960\" target=\"_blank\">https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960</a>. </p>\n<p>The information matching antibody names to ensemble gene IDs that you're looking for is in the attached Excel file.</p>\n<p>Also the <code>Mouse-</code> and <code>Rat-</code> features are \"isogenic controls\" that should not bind human proteins. You can think of these as negative controls measuring the background signal in the CITE ADT measurements. We thought it could be interesting to include these to see if is a way competitors can predict the baseline ADT based on the RNA data.</p>\n<p>EDIT 2022-10-12: New Excel from BioLegend with compete Ensembl ID Information: <code>TotalSeq_B_Universal_Cocktail_v1_Antibodies_399904_Barcodes_BioLegendUpdate.xlsx</code></p>",
      "rawMarkdown": "You can find full information about the antibody panel for the CITE here: https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960. \n\nThe information matching antibody names to ensemble gene IDs that you're looking for is in the attached Excel file.\n\nAlso the `Mouse-` and `Rat-` features are \"isogenic controls\" that should not bind human proteins. You can think of these as negative controls measuring the background signal in the CITE ADT measurements. We thought it could be interesting to include these to see if is a way competitors can predict the baseline ADT based on the RNA data.\n\nEDIT 2022-10-12: New Excel from BioLegend with compete Ensembl ID Information: `TotalSeq_B_Universal_Cocktail_v1_Antibodies_399904_Barcodes_BioLegendUpdate.xlsx`",
      "votes": 6,
      "replies": [
        {
          "id": 1983696,
          "postDate": "2022-10-12T07:26:51.467Z",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Thanks for the file ! <br>\nBut it is slightly broken - it seems it is not your fault - but it is on vendor side like that.<br>\nMay be you can inform them</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F5bb78d93f513ac6b0f24d66ffa7568a7%2Fphoto_2022-10-12_09-25-10.jpg?generation=1665559549382372&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@danielburkhardt Thanks for the file ! \nBut it is slightly broken - it seems it is not your fault - but it is on vendor side like that.\nMay be you can inform them\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F5bb78d93f513ac6b0f24d66ffa7568a7%2Fphoto_2022-10-12_09-25-10.jpg?generation=1665559549382372&alt=media)\n\n\n"
        },
        {
          "id": 1983807,
          "postDate": "2022-10-12T08:47:23.117Z",
          "content": "<p>Correspondence (with some gaps) can be also found here:<br>\n<a href=\"https://www.kaggle.com/code/geraseva/magic-parameters?scriptVersionId=107776618&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/geraseva/magic-parameters?scriptVersionId=107776618&amp;cellId=21</a><br>\nThanks to Liza Geraseva</p>",
          "rawMarkdown": "Correspondence (with some gaps) can be also found here:\nhttps://www.kaggle.com/code/geraseva/magic-parameters?scriptVersionId=107776618&cellId=21\nThanks to Liza Geraseva"
        },
        {
          "id": 1984553,
          "postDate": "2022-10-12T17:53:37.840Z",
          "content": "<p>I didn't notice this error! I just sent an email to BioLegend support, they obviously made a mistake.</p>\n<p>I can't be certain about this, but I think these are the correct Ensemble IDs for those columns based on manually searching the Ensembl database:</p>\n<table>\n<thead>\n<tr>\n<th>DNA_ID</th>\n<th>Description</th>\n<th>clone</th>\n<th>Sequence</th>\n<th>Ensemble ID</th>\n<th>Gene symbol</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0174</td>\n<td>anti-human CD58   (LFA-3)</td>\n<td>TS2/9</td>\n<td>GTTCCTATGGACGAC</td>\n<td>ENSG00000116815</td>\n<td>CD58</td>\n</tr>\n<tr>\n<td>B0176</td>\n<td>anti-human CD39</td>\n<td>A1</td>\n<td>TTACCTGGTATCCGT</td>\n<td>ENSG00000138185</td>\n<td>ENTPD1</td>\n</tr>\n<tr>\n<td>B0179</td>\n<td>anti-human CX3CR1</td>\n<td>K0124E1</td>\n<td>AGTATCGTCTCTGGG</td>\n<td>ENSG00000168329</td>\n<td>CX3CR1</td>\n</tr>\n<tr>\n<td>B0180</td>\n<td>anti-human CD24</td>\n<td>ML5</td>\n<td>AGATTCCTTCGTGTT</td>\n<td>ENSG00000272398</td>\n<td>CD24</td>\n</tr>\n<tr>\n<td>B0181</td>\n<td>anti-human CD21</td>\n<td>Bu32</td>\n<td>AACCTAGTAGTTCGG</td>\n<td>ENSG00000117322</td>\n<td>CR2</td>\n</tr>\n<tr>\n<td>B0185</td>\n<td>anti-human CD11a</td>\n<td>TS2/4</td>\n<td>TATATCCTTGTGAGC</td>\n<td>ENSG00000005844</td>\n<td>ITGAL</td>\n</tr>\n<tr>\n<td>B0187</td>\n<td>anti-human CD79b   (Igβ)</td>\n<td>CB3-1</td>\n<td>ATTCTTCAACCGAAG</td>\n<td>ENSG00000007312</td>\n<td>CD79B</td>\n</tr>\n<tr>\n<td>B0189</td>\n<td>anti-human CD244   (2B4)</td>\n<td>C1.7</td>\n<td>TCGCTTGGATGGTAG</td>\n<td>ENSG00000122223</td>\n<td>CD244</td>\n</tr>\n<tr>\n<td>B0206</td>\n<td>anti-human CD169</td>\n<td>7-239</td>\n<td>TACTCAGCGTGTTTG</td>\n<td>ENSG00000088827</td>\n<td>SIGLEC1</td>\n</tr>\n<tr>\n<td>B0214</td>\n<td>anti-human/mouse   integrin β7</td>\n<td>FIB504</td>\n<td>TCCTTGGATGTACCG</td>\n<td>ENSG00000139626</td>\n<td>ITGB7</td>\n</tr>\n<tr>\n<td>B0215</td>\n<td>anti-human CD268   (BAFF-R)</td>\n<td>11C1</td>\n<td>CGAAGTCGATCCGTA</td>\n<td>ENSG00000159958</td>\n<td>TNFRSF13C</td>\n</tr>\n<tr>\n<td>B0216</td>\n<td>anti-human CD42b</td>\n<td>HIP1</td>\n<td>TCCTAGTACCGAAGT</td>\n<td>ENSG00000185245</td>\n<td>GP1BA</td>\n</tr>\n<tr>\n<td>B0217</td>\n<td>anti-human CD54</td>\n<td>HA58</td>\n<td>CTGATAGACTTGAGT</td>\n<td>ENSG00000090339</td>\n<td>ICAM1</td>\n</tr>\n<tr>\n<td>B0218</td>\n<td>anti-human CD62P   (P-Selectin)</td>\n<td>AK4</td>\n<td>CCTTCCGTATCCCTT</td>\n<td>ENSG00000174175</td>\n<td>SELP</td>\n</tr>\n<tr>\n<td>B0219</td>\n<td>anti-human CD119   (IFN-γ R α chain)</td>\n<td>GIR-208</td>\n<td>TGTGTATTCCCTTGT</td>\n<td>ENSG00000027697</td>\n<td>IFNGR1</td>\n</tr>\n<tr>\n<td>B0224</td>\n<td>anti-human TCR α/β</td>\n<td>IP26</td>\n<td>CGTAACGTAGAGCGA</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n</tbody>\n</table>\n<p>The last one, <code>anti-human TCR α/β</code> targets a multi-gene protein complex so there's not a single Ensembl ID for that. This page has more specific information about the antibody clone and the antigen it targets.</p>\n<p>I'll update you when I heard back from BioLegend!</p>\n<p>EDIT: Attached is the updated Excel sheet BioLegend just sent me.</p>",
          "rawMarkdown": "I didn't notice this error! I just sent an email to BioLegend support, they obviously made a mistake.\n\nI can't be certain about this, but I think these are the correct Ensemble IDs for those columns based on manually searching the Ensembl database:\n\n| DNA_ID |              Description             |  clone  |     Sequence    | Ensemble ID     | Gene symbol |\n|:------:|:------------------------------------:|:-------:|:---------------:|-----------------|-------------|\n|  B0174 | anti-human CD58   (LFA-3)            | TS2/9   | GTTCCTATGGACGAC | ENSG00000116815 | CD58        |\n|  B0176 | anti-human CD39                      | A1      | TTACCTGGTATCCGT | ENSG00000138185 | ENTPD1      |\n|  B0179 | anti-human CX3CR1                    | K0124E1 | AGTATCGTCTCTGGG | ENSG00000168329 | CX3CR1      |\n|  B0180 | anti-human CD24                      | ML5     | AGATTCCTTCGTGTT | ENSG00000272398 | CD24        |\n|  B0181 | anti-human CD21                      | Bu32    | AACCTAGTAGTTCGG | ENSG00000117322 | CR2         |\n|  B0185 | anti-human CD11a                     | TS2/4   | TATATCCTTGTGAGC | ENSG00000005844 | ITGAL       |\n|  B0187 | anti-human CD79b   (Igβ)             | CB3-1   | ATTCTTCAACCGAAG | ENSG00000007312 | CD79B       |\n|  B0189 | anti-human CD244   (2B4)             | C1.7    | TCGCTTGGATGGTAG | ENSG00000122223 | CD244       |\n|  B0206 | anti-human CD169                     | 7-239   | TACTCAGCGTGTTTG | ENSG00000088827 | SIGLEC1     |\n|  B0214 | anti-human/mouse   integrin β7       | FIB504  | TCCTTGGATGTACCG | ENSG00000139626 | ITGB7       |\n|  B0215 | anti-human CD268   (BAFF-R)          | 11C1    | CGAAGTCGATCCGTA | ENSG00000159958 | TNFRSF13C   |\n|  B0216 | anti-human CD42b                     | HIP1    | TCCTAGTACCGAAGT | ENSG00000185245 | GP1BA       |\n|  B0217 | anti-human CD54                      | HA58    | CTGATAGACTTGAGT | ENSG00000090339 | ICAM1       |\n|  B0218 | anti-human CD62P   (P-Selectin)      | AK4     | CCTTCCGTATCCCTT | ENSG00000174175 | SELP        |\n|  B0219 | anti-human CD119   (IFN-γ R α chain) | GIR-208 | TGTGTATTCCCTTGT | ENSG00000027697 | IFNGR1      |\n|  B0224 | anti-human TCR α/β                   | IP26    | CGTAACGTAGAGCGA | NaN             | NaN         |\n\n\nThe last one, `anti-human TCR α/β` targets a multi-gene protein complex so there's not a single Ensembl ID for that. This page has more specific information about the antibody clone and the antigen it targets.\n\nI'll update you when I heard back from BioLegend!\n\nEDIT: Attached is the updated Excel sheet BioLegend just sent me.",
          "votes": 1,
          "replies": [
            {
              "id": 2071060,
              "postDate": "2022-12-20T15:16:51.060Z",
              "content": "<p>The notebook:<br>\n<a href=\"https://www.kaggle.com/alexandervc/cd-names-to-ensembl-symbol-names-correspondence\" target=\"_blank\">https://www.kaggle.com/alexandervc/cd-names-to-ensembl-symbol-names-correspondence</a><br>\nAnd the output file:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/cd-names-to-ensembl-symbol-names-correspondence/data?scriptVersionId=114325820&amp;select=CD_to_EnsemblSymbol_correspondence_NIPS2022.csv\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cd-names-to-ensembl-symbol-names-correspondence/data?scriptVersionId=114325820&amp;select=CD_to_EnsemblSymbol_correspondence_NIPS2022.csv</a></p>\n<p>Explains some caveats of the correspondence CD-&gt;RNA in provided Excel file.<br>\nAnd generates more easy to use CSV correspondence file.</p>",
              "rawMarkdown": "The notebook:\nhttps://www.kaggle.com/alexandervc/cd-names-to-ensembl-symbol-names-correspondence\nAnd the output file:\nhttps://www.kaggle.com/code/alexandervc/cd-names-to-ensembl-symbol-names-correspondence/data?scriptVersionId=114325820&select=CD_to_EnsemblSymbol_correspondence_NIPS2022.csv\n\nExplains some caveats of the correspondence CD->RNA in provided Excel file.\nAnd generates more easy to use CSV correspondence file.\n"
            }
          ]
        }
      ]
    },
    {
      "id": 1952125,
      "postDate": "2022-09-23T13:12:30.767Z",
      "content": "<p>Would you be so kind to comment what these means : <br>\n'HLA-A-B-C', 'TIGIT', 'Mouse-IgG1', 'Mouse-IgG2a', 'Mouse-IgG2b',<br>\n   'Rat-IgG2b', 'Podoplanin', 'IgM', 'KLRG1', 'HLA-DR', 'CX3CR1',<br>\n   'integrinB7', 'TCR', 'Rat-IgG1', 'Rat-IgG2a', 'FceRIa', 'IgD',<br>\n   'TCRVa7.2', 'TCRVd2', 'LOX-1', 'HLA-E'</p>\n<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
      "rawMarkdown": "Would you be so kind to comment what these means : \n'HLA-A-B-C', 'TIGIT', 'Mouse-IgG1', 'Mouse-IgG2a', 'Mouse-IgG2b',\n   'Rat-IgG2b', 'Podoplanin', 'IgM', 'KLRG1', 'HLA-DR', 'CX3CR1',\n   'integrinB7', 'TCR', 'Rat-IgG1', 'Rat-IgG2a', 'FceRIa', 'IgD',\n   'TCRVa7.2', 'TCRVd2', 'LOX-1', 'HLA-E'\n\n@danielburkhardt ",
      "votes": 6
    },
    {
      "id": 1956786,
      "postDate": "2022-09-26T15:42:22.447Z",
      "content": "<p>Idea: <br>\nThe possible idea to use it - predict not them, but a their average.<br>\nThen put them all equal to that average.  One may subtract that average from all predictions (to make them zero) - or may not - it does not matter  because metric - is correlation coefficient which does not depend on adding constant. </p>\n<p>If they are - just noise, than  it should improve the prediction. </p>",
      "rawMarkdown": "Idea: \nThe possible idea to use it - predict not them, but a their average.\nThen put them all equal to that average.  One may subtract that average from all predictions (to make them zero) - or may not - it does not matter  because metric - is correlation coefficient which does not depend on adding constant. \n \nIf they are - just noise, than  it should improve the prediction. ",
      "votes": 1,
      "replies": [
        {
          "id": 1961322,
          "postDate": "2022-09-29T05:49:29.873Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1961472,
          "postDate": "2022-09-29T07:40:03.117Z",
          "content": "<p>No.<br>\nPlease read the answer from orgs.</p>",
          "rawMarkdown": "No.\nPlease read the answer from orgs.\n"
        }
      ]
    },
    {
      "id": 1956916,
      "postDate": "2022-09-26T17:01:50.570Z",
      "content": "<p>These are all differnent protein names, for pure prediction, I think you do not need to know too much about their biological functions. If you want to know more, I recommend to search in NCBI database</p>",
      "rawMarkdown": "These are all differnent protein names, for pure prediction, I think you do not need to know too much about their biological functions. If you want to know more, I recommend to search in NCBI database",
      "replies": [
        {
          "id": 1962629,
          "postDate": "2022-09-29T19:59:27.030Z",
          "content": "<p>Or you can get info by some packages<br>\nlike Python \"mygene\":<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-genes-eda\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-genes-eda</a></p>",
          "rawMarkdown": "Or you can get info by some packages\nlike Python \"mygene\":\nhttps://www.kaggle.com/code/alexandervc/mmscel-genes-eda\n"
        },
        {
          "id": 1970003,
          "postDate": "2022-10-03T19:17:55.847Z",
          "content": "<p>thanks for your sharing</p>",
          "rawMarkdown": "thanks for your sharing"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1952448,
      "author_name": "Daniel Burkhardt",
      "author_url": "",
      "post_date": "2022-09-23T17:14:02.977000",
      "content": "<p>You can find full information about the antibody panel for the CITE here: <a href=\"https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960\" target=\"_blank\">https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960</a>. </p>\n<p>The information matching antibody names to ensemble gene IDs that you're looking for is in the attached Excel file.</p>\n<p>Also the <code>Mouse-</code> and <code>Rat-</code> features are \"isogenic controls\" that should not bind human proteins. You can think of these as negative controls measuring the background signal in the CITE ADT measurements. We thought it could be interesting to include these to see if is a way competitors can predict the baseline ADT based on the RNA data.</p>\n<p>EDIT 2022-10-12: New Excel from BioLegend with compete Ensembl ID Information: <code>TotalSeq_B_Universal_Cocktail_v1_Antibodies_399904_Barcodes_BioLegendUpdate.xlsx</code></p>",
      "votes": 6,
      "replies": [
        {
          "id": 1983696,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-10-12T07:26:51.467000",
          "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Thanks for the file ! <br>\nBut it is slightly broken - it seems it is not your fault - but it is on vendor side like that.<br>\nMay be you can inform them</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F5bb78d93f513ac6b0f24d66ffa7568a7%2Fphoto_2022-10-12_09-25-10.jpg?generation=1665559549382372&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1983807,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-10-12T08:47:23.117000",
          "content": "<p>Correspondence (with some gaps) can be also found here:<br>\n<a href=\"https://www.kaggle.com/code/geraseva/magic-parameters?scriptVersionId=107776618&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/geraseva/magic-parameters?scriptVersionId=107776618&amp;cellId=21</a><br>\nThanks to Liza Geraseva</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1984553,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-10-12T17:53:37.840000",
          "content": "<p>I didn't notice this error! I just sent an email to BioLegend support, they obviously made a mistake.</p>\n<p>I can't be certain about this, but I think these are the correct Ensemble IDs for those columns based on manually searching the Ensembl database:</p>\n<table>\n<thead>\n<tr>\n<th>DNA_ID</th>\n<th>Description</th>\n<th>clone</th>\n<th>Sequence</th>\n<th>Ensemble ID</th>\n<th>Gene symbol</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0174</td>\n<td>anti-human CD58   (LFA-3)</td>\n<td>TS2/9</td>\n<td>GTTCCTATGGACGAC</td>\n<td>ENSG00000116815</td>\n<td>CD58</td>\n</tr>\n<tr>\n<td>B0176</td>\n<td>anti-human CD39</td>\n<td>A1</td>\n<td>TTACCTGGTATCCGT</td>\n<td>ENSG00000138185</td>\n<td>ENTPD1</td>\n</tr>\n<tr>\n<td>B0179</td>\n<td>anti-human CX3CR1</td>\n<td>K0124E1</td>\n<td>AGTATCGTCTCTGGG</td>\n<td>ENSG00000168329</td>\n<td>CX3CR1</td>\n</tr>\n<tr>\n<td>B0180</td>\n<td>anti-human CD24</td>\n<td>ML5</td>\n<td>AGATTCCTTCGTGTT</td>\n<td>ENSG00000272398</td>\n<td>CD24</td>\n</tr>\n<tr>\n<td>B0181</td>\n<td>anti-human CD21</td>\n<td>Bu32</td>\n<td>AACCTAGTAGTTCGG</td>\n<td>ENSG00000117322</td>\n<td>CR2</td>\n</tr>\n<tr>\n<td>B0185</td>\n<td>anti-human CD11a</td>\n<td>TS2/4</td>\n<td>TATATCCTTGTGAGC</td>\n<td>ENSG00000005844</td>\n<td>ITGAL</td>\n</tr>\n<tr>\n<td>B0187</td>\n<td>anti-human CD79b   (Igβ)</td>\n<td>CB3-1</td>\n<td>ATTCTTCAACCGAAG</td>\n<td>ENSG00000007312</td>\n<td>CD79B</td>\n</tr>\n<tr>\n<td>B0189</td>\n<td>anti-human CD244   (2B4)</td>\n<td>C1.7</td>\n<td>TCGCTTGGATGGTAG</td>\n<td>ENSG00000122223</td>\n<td>CD244</td>\n</tr>\n<tr>\n<td>B0206</td>\n<td>anti-human CD169</td>\n<td>7-239</td>\n<td>TACTCAGCGTGTTTG</td>\n<td>ENSG00000088827</td>\n<td>SIGLEC1</td>\n</tr>\n<tr>\n<td>B0214</td>\n<td>anti-human/mouse   integrin β7</td>\n<td>FIB504</td>\n<td>TCCTTGGATGTACCG</td>\n<td>ENSG00000139626</td>\n<td>ITGB7</td>\n</tr>\n<tr>\n<td>B0215</td>\n<td>anti-human CD268   (BAFF-R)</td>\n<td>11C1</td>\n<td>CGAAGTCGATCCGTA</td>\n<td>ENSG00000159958</td>\n<td>TNFRSF13C</td>\n</tr>\n<tr>\n<td>B0216</td>\n<td>anti-human CD42b</td>\n<td>HIP1</td>\n<td>TCCTAGTACCGAAGT</td>\n<td>ENSG00000185245</td>\n<td>GP1BA</td>\n</tr>\n<tr>\n<td>B0217</td>\n<td>anti-human CD54</td>\n<td>HA58</td>\n<td>CTGATAGACTTGAGT</td>\n<td>ENSG00000090339</td>\n<td>ICAM1</td>\n</tr>\n<tr>\n<td>B0218</td>\n<td>anti-human CD62P   (P-Selectin)</td>\n<td>AK4</td>\n<td>CCTTCCGTATCCCTT</td>\n<td>ENSG00000174175</td>\n<td>SELP</td>\n</tr>\n<tr>\n<td>B0219</td>\n<td>anti-human CD119   (IFN-γ R α chain)</td>\n<td>GIR-208</td>\n<td>TGTGTATTCCCTTGT</td>\n<td>ENSG00000027697</td>\n<td>IFNGR1</td>\n</tr>\n<tr>\n<td>B0224</td>\n<td>anti-human TCR α/β</td>\n<td>IP26</td>\n<td>CGTAACGTAGAGCGA</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n</tbody>\n</table>\n<p>The last one, <code>anti-human TCR α/β</code> targets a multi-gene protein complex so there's not a single Ensembl ID for that. This page has more specific information about the antibody clone and the antigen it targets.</p>\n<p>I'll update you when I heard back from BioLegend!</p>\n<p>EDIT: Attached is the updated Excel sheet BioLegend just sent me.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2071060,
              "author_name": "Alexander Chervov",
              "author_url": "",
              "post_date": "2022-12-20T15:16:51.060000",
              "content": "<p>The notebook:<br>\n<a href=\"https://www.kaggle.com/alexandervc/cd-names-to-ensembl-symbol-names-correspondence\" target=\"_blank\">https://www.kaggle.com/alexandervc/cd-names-to-ensembl-symbol-names-correspondence</a><br>\nAnd the output file:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/cd-names-to-ensembl-symbol-names-correspondence/data?scriptVersionId=114325820&amp;select=CD_to_EnsemblSymbol_correspondence_NIPS2022.csv\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cd-names-to-ensembl-symbol-names-correspondence/data?scriptVersionId=114325820&amp;select=CD_to_EnsemblSymbol_correspondence_NIPS2022.csv</a></p>\n<p>Explains some caveats of the correspondence CD-&gt;RNA in provided Excel file.<br>\nAnd generates more easy to use CSV correspondence file.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1956786,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-09-26T15:42:22.447000",
      "content": "<p>Idea: <br>\nThe possible idea to use it - predict not them, but a their average.<br>\nThen put them all equal to that average.  One may subtract that average from all predictions (to make them zero) - or may not - it does not matter  because metric - is correlation coefficient which does not depend on adding constant. </p>\n<p>If they are - just noise, than  it should improve the prediction. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1961322,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-29T05:49:29.873000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1961472,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-09-29T07:40:03.117000",
          "content": "<p>No.<br>\nPlease read the answer from orgs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1956916,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "2022-09-26T17:01:50.570000",
      "content": "<p>These are all differnent protein names, for pure prediction, I think you do not need to know too much about their biological functions. If you want to know more, I recommend to search in NCBI database</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1962629,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-09-29T19:59:27.030000",
          "content": "<p>Or you can get info by some packages<br>\nlike Python \"mygene\":<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-genes-eda\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-genes-eda</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1970003,
          "author_name": "jinyang18",
          "author_url": "",
          "post_date": "2022-10-03T19:17:55.847000",
          "content": "<p>thanks for your sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1952448": "You can find full information about the antibody panel for the CITE here: https://www.biolegend.com/de-at/products/totalseq-b-human-universal-cocktail-v1dot0-20960. \n\nThe information matching antibody names to ensemble gene IDs that you're looking for is in the attached Excel file.\n\nAlso the `Mouse-` and `Rat-` features are \"isogenic controls\" that should not bind human proteins. You can think of these as negative controls measuring the background signal in the CITE ADT measurements. We thought it could be interesting to include these to see if is a way competitors can predict the baseline ADT based on the RNA data.\n\nEDIT 2022-10-12: New Excel from BioLegend with compete Ensembl ID Information: `TotalSeq_B_Universal_Cocktail_v1_Antibodies_399904_Barcodes_BioLegendUpdate.xlsx`",
    "1952125": "Would you be so kind to comment what these means : \n'HLA-A-B-C', 'TIGIT', 'Mouse-IgG1', 'Mouse-IgG2a', 'Mouse-IgG2b',\n   'Rat-IgG2b', 'Podoplanin', 'IgM', 'KLRG1', 'HLA-DR', 'CX3CR1',\n   'integrinB7', 'TCR', 'Rat-IgG1', 'Rat-IgG2a', 'FceRIa', 'IgD',\n   'TCRVa7.2', 'TCRVd2', 'LOX-1', 'HLA-E'\n\n@danielburkhardt ",
    "1956786": "Idea: \nThe possible idea to use it - predict not them, but a their average.\nThen put them all equal to that average.  One may subtract that average from all predictions (to make them zero) - or may not - it does not matter  because metric - is correlation coefficient which does not depend on adding constant. \n \nIf they are - just noise, than  it should improve the prediction. ",
    "1956916": "These are all differnent protein names, for pure prediction, I think you do not need to know too much about their biological functions. If you want to know more, I recommend to search in NCBI database"
  }
}