{
  "id": 221025,
  "title": "Silly Question",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/221025",
  "author_name": "chemdatafarmer",
  "post_date": "2021-02-20T16:28:27.564000",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I'm having trouble with the simplest thing…I'm trying to take a dataframe that holds all the test data (df), and isolate all the rows where only 1 label is present in the image.</p>\n<p>To do this I tried the following line of code.</p>\n<p>test = df[~df['Label'].str.contains('|')]</p>\n<p>However, when I run the above line, I just get an empty dataframe even though I know there are examples where the Label column has single values without '|' in them. When I do the following, I simply get a copy of the original dataframe.</p>\n<p>test_2 = df[df['Label'].str.contains('|')]</p>\n<p>I know this is pretty basic, but I'm here to learn so I figured I'd ask…Any help would be greatly appreciated. I'm running this code on a Kaggle notebook.</p>",
  "messages": [
    {
      "id": 1211920,
      "postDate": "2021-02-20T17:02:44.133Z",
      "content": "<p>you can try</p>\n<pre>df[~df['Label'].str.contains('\\|')]\n</pre>\n<p><code>|</code> has a special meaning <code>OR</code>.</p>\n<p>ex.  <code>3|9</code> means 3 OR 9.</p>\n<pre>df[df['Label'].str.contains('3|9')]\n7    5c68183e-bb99-11e8-b2b9-ac1f6b6435d0    13|0\n11    5f1af6b4-bb99-11e8-b2b9-ac1f6b6435d0    3\n13    5fb9edb4-bb99-11e8-b2b9-ac1f6b6435d0    9\n17    636e164c-bb99-11e8-b2b9-ac1f6b6435d0    3\n18    5b99d3e8-bb99-11e8-b2b9-ac1f6b6435d0    13\n</pre>",
      "rawMarkdown": "you can try\n<pre>\ndf[~df['Label'].str.contains('\\|')]\n</pre>\n`|` has a special meaning `OR`.\n\nex.  `3|9` means 3 OR 9.\n<pre>\ndf[df['Label'].str.contains('3|9')]\n7\t5c68183e-bb99-11e8-b2b9-ac1f6b6435d0\t13|0\n11\t5f1af6b4-bb99-11e8-b2b9-ac1f6b6435d0\t3\n13\t5fb9edb4-bb99-11e8-b2b9-ac1f6b6435d0\t9\n17\t636e164c-bb99-11e8-b2b9-ac1f6b6435d0\t3\n18\t5b99d3e8-bb99-11e8-b2b9-ac1f6b6435d0\t13\n</pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1211945,
          "postDate": "2021-02-20T17:26:33.797Z",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> ! This solved my problem and your description was quite helpful.</p>",
          "rawMarkdown": "Thank you very much @its7171 ! This solved my problem and your description was quite helpful.",
          "votes": 1
        },
        {
          "id": 1212161,
          "postDate": "2021-02-21T00:13:21.860Z",
          "content": "<p>I had the same problem exactly previously. This is the solution I ended up at as well.</p>",
          "rawMarkdown": "I had the same problem exactly previously. This is the solution I ended up at as well.",
          "votes": 1
        },
        {
          "id": 1212254,
          "postDate": "2021-02-21T03:46:13.467Z",
          "content": "<p>Funny how one character can make all the difference haha</p>",
          "rawMarkdown": "Funny how one character can make all the difference haha"
        }
      ]
    },
    {
      "id": 1211889,
      "postDate": "2021-02-20T16:28:27.563Z",
      "content": "<p>Hi All,</p>\n<p>I'm having trouble with the simplest thing…I'm trying to take a dataframe that holds all the test data (df), and isolate all the rows where only 1 label is present in the image.</p>\n<p>To do this I tried the following line of code.</p>\n<p>test = df[~df['Label'].str.contains('|')]</p>\n<p>However, when I run the above line, I just get an empty dataframe even though I know there are examples where the Label column has single values without '|' in them. When I do the following, I simply get a copy of the original dataframe.</p>\n<p>test_2 = df[df['Label'].str.contains('|')]</p>\n<p>I know this is pretty basic, but I'm here to learn so I figured I'd ask…Any help would be greatly appreciated. I'm running this code on a Kaggle notebook.</p>",
      "rawMarkdown": "Hi All,\n\nI'm having trouble with the simplest thing...I'm trying to take a dataframe that holds all the test data (df), and isolate all the rows where only 1 label is present in the image.\n\nTo do this I tried the following line of code.\n\ntest = df[~df['Label'].str.contains('|')]\n\nHowever, when I run the above line, I just get an empty dataframe even though I know there are examples where the Label column has single values without '|' in them. When I do the following, I simply get a copy of the original dataframe.\n\ntest_2 = df[df['Label'].str.contains('|')]\n\nI know this is pretty basic, but I'm here to learn so I figured I'd ask...Any help would be greatly appreciated. I'm running this code on a Kaggle notebook.",
      "votes": 2
    },
    {
      "id": 1212366,
      "postDate": "2021-02-21T06:48:04.127Z",
      "content": "<p>You will miss images if you only filter by '|' because of duplicate labels ('9|9', for example, appears 6 times):</p>\n<p><code>df['Label_lst'] = df['Label'].str.split('|')</code><br>\n<code>df[df['Label'] == '9|9']</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6954</td>\n<td>abdff2be-bba9-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>12535</td>\n<td>0748041e-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>12691</td>\n<td>5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<p>You can solve this with:</p>\n<p><code>df['Label_lst'] = df['Label'].str.split('|')</code><br>\n<code>df['Label_lst'] = df['Label_lst'].map(pd.unique)</code><br>\n<code>df[df['Label'] == '9|9']</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6954</td>\n<td>abdff2be-bba9-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>12535</td>\n<td>0748041e-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>12691</td>\n<td>5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<p>Filter images with 1 label:</p>\n<p><code>df[df['Label_lst'].map(len) == 1]</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5</td>\n<td>5e22a522-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>0</td>\n<td>[0]</td>\n</tr>\n<tr>\n<td>6</td>\n<td>5f79a114-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>14</td>\n<td>[14]</td>\n</tr>\n<tr>\n<td>9</td>\n<td>5c801c04-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>14</td>\n<td>[14]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "You will miss images if you only filter by '|' because of duplicate labels ('9\\|9', for example, appears 6 times):\n\n`df['Label_lst'] = df['Label'].str.split('|')`\n`df[df['Label'] == '9|9']`\n\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|6954\t | abdff2be-bba9-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|12535\t| 0748041e-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|12691\t| 5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|...| ...\t|... |... |\n\n\nYou can solve this with:\n\n`df['Label_lst'] = df['Label'].str.split('|')`\n`df['Label_lst'] = df['Label_lst'].map(pd.unique)`\n`df[df['Label'] == '9|9']`\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|6954\t | abdff2be-bba9-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|12535\t| 0748041e-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|12691\t| 5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|...| ...\t|... |... |\n\nFilter images with 1 label:\n\n`df[df['Label_lst'].map(len) == 1]`\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|5| 5e22a522-bb99-11e8-b2b9-ac1f6b6435d0\t|0 |[0] |\n|6| 5f79a114-bb99-11e8-b2b9-ac1f6b6435d0\t|14 |[14] |\n|9| 5c801c04-bb99-11e8-b2b9-ac1f6b6435d0\t|14 |[14] |\n|...| ...\t|... |... |\n",
      "replies": [
        {
          "id": 1212529,
          "postDate": "2021-02-21T10:00:11.713Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1211892,
      "postDate": "2021-02-20T16:30:41.527Z",
      "content": "<p>sorry, I missed a bracket in my explanation. The problem I'm having still holds. test = df[~df['Label'].str.contains('|')] just gives me an empty dataframe…</p>",
      "rawMarkdown": "sorry, I missed a bracket in my explanation. The problem I'm having still holds. test = df[~df['Label'].str.contains('|')] just gives me an empty dataframe..."
    },
    {
      "id": 1212133,
      "postDate": "2021-02-20T22:53:59.400Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1212143,
          "postDate": "2021-02-20T23:16:14.400Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">@weka511</a> thanks for your input! I'm also using Python, the dataframe is a Pandas dataframe. I probably should have mentioned that. I think the trouble I ran into was it was recognizing '|' as an OR operator (see comment from tito). When I escaped this special meaning with '|' all worked just fine.  </p>",
          "rawMarkdown": "Hi @weka511 thanks for your input! I'm also using Python, the dataframe is a Pandas dataframe. I probably should have mentioned that. I think the trouble I ran into was it was recognizing '|' as an OR operator (see comment from tito). When I escaped this special meaning with '\\|' all worked just fine.  "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1211920,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2021-02-20T17:02:44.133000",
      "content": "<p>you can try</p>\n<pre>df[~df['Label'].str.contains('\\|')]\n</pre>\n<p><code>|</code> has a special meaning <code>OR</code>.</p>\n<p>ex.  <code>3|9</code> means 3 OR 9.</p>\n<pre>df[df['Label'].str.contains('3|9')]\n7    5c68183e-bb99-11e8-b2b9-ac1f6b6435d0    13|0\n11    5f1af6b4-bb99-11e8-b2b9-ac1f6b6435d0    3\n13    5fb9edb4-bb99-11e8-b2b9-ac1f6b6435d0    9\n17    636e164c-bb99-11e8-b2b9-ac1f6b6435d0    3\n18    5b99d3e8-bb99-11e8-b2b9-ac1f6b6435d0    13\n</pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1211945,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2021-02-20T17:26:33.797000",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> ! This solved my problem and your description was quite helpful.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212161,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-21T00:13:21.860000",
          "content": "<p>I had the same problem exactly previously. This is the solution I ended up at as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212254,
          "author_name": "Arka Saha",
          "author_url": "",
          "post_date": "2021-02-21T03:46:13.467000",
          "content": "<p>Funny how one character can make all the difference haha</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1212366,
      "author_name": "Dan Presil",
      "author_url": "",
      "post_date": "2021-02-21T06:48:04.127000",
      "content": "<p>You will miss images if you only filter by '|' because of duplicate labels ('9|9', for example, appears 6 times):</p>\n<p><code>df['Label_lst'] = df['Label'].str.split('|')</code><br>\n<code>df[df['Label'] == '9|9']</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6954</td>\n<td>abdff2be-bba9-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>12535</td>\n<td>0748041e-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>12691</td>\n<td>5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9, 9]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<p>You can solve this with:</p>\n<p><code>df['Label_lst'] = df['Label'].str.split('|')</code><br>\n<code>df['Label_lst'] = df['Label_lst'].map(pd.unique)</code><br>\n<code>df[df['Label'] == '9|9']</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6954</td>\n<td>abdff2be-bba9-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>12535</td>\n<td>0748041e-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>12691</td>\n<td>5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0</td>\n<td>9|9</td>\n<td>[9]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<p>Filter images with 1 label:</p>\n<p><code>df[df['Label_lst'].map(len) == 1]</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ID</th>\n<th>Label</th>\n<th>Label_lst</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5</td>\n<td>5e22a522-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>0</td>\n<td>[0]</td>\n</tr>\n<tr>\n<td>6</td>\n<td>5f79a114-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>14</td>\n<td>[14]</td>\n</tr>\n<tr>\n<td>9</td>\n<td>5c801c04-bb99-11e8-b2b9-ac1f6b6435d0</td>\n<td>14</td>\n<td>[14]</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>",
      "votes": 0,
      "replies": [
        {
          "id": 1212529,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-21T10:00:11.713000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1211892,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2021-02-20T16:30:41.527000",
      "content": "<p>sorry, I missed a bracket in my explanation. The problem I'm having still holds. test = df[~df['Label'].str.contains('|')] just gives me an empty dataframe…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1212133,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-20T22:53:59.400000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1212143,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2021-02-20T23:16:14.400000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">@weka511</a> thanks for your input! I'm also using Python, the dataframe is a Pandas dataframe. I probably should have mentioned that. I think the trouble I ran into was it was recognizing '|' as an OR operator (see comment from tito). When I escaped this special meaning with '|' all worked just fine.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1211920": "you can try\n<pre>\ndf[~df['Label'].str.contains('\\|')]\n</pre>\n`|` has a special meaning `OR`.\n\nex.  `3|9` means 3 OR 9.\n<pre>\ndf[df['Label'].str.contains('3|9')]\n7\t5c68183e-bb99-11e8-b2b9-ac1f6b6435d0\t13|0\n11\t5f1af6b4-bb99-11e8-b2b9-ac1f6b6435d0\t3\n13\t5fb9edb4-bb99-11e8-b2b9-ac1f6b6435d0\t9\n17\t636e164c-bb99-11e8-b2b9-ac1f6b6435d0\t3\n18\t5b99d3e8-bb99-11e8-b2b9-ac1f6b6435d0\t13\n</pre>",
    "1211889": "Hi All,\n\nI'm having trouble with the simplest thing...I'm trying to take a dataframe that holds all the test data (df), and isolate all the rows where only 1 label is present in the image.\n\nTo do this I tried the following line of code.\n\ntest = df[~df['Label'].str.contains('|')]\n\nHowever, when I run the above line, I just get an empty dataframe even though I know there are examples where the Label column has single values without '|' in them. When I do the following, I simply get a copy of the original dataframe.\n\ntest_2 = df[df['Label'].str.contains('|')]\n\nI know this is pretty basic, but I'm here to learn so I figured I'd ask...Any help would be greatly appreciated. I'm running this code on a Kaggle notebook.",
    "1212366": "You will miss images if you only filter by '|' because of duplicate labels ('9\\|9', for example, appears 6 times):\n\n`df['Label_lst'] = df['Label'].str.split('|')`\n`df[df['Label'] == '9|9']`\n\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|6954\t | abdff2be-bba9-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|12535\t| 0748041e-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|12691\t| 5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9, 9] |\n|...| ...\t|... |... |\n\n\nYou can solve this with:\n\n`df['Label_lst'] = df['Label'].str.split('|')`\n`df['Label_lst'] = df['Label_lst'].map(pd.unique)`\n`df[df['Label'] == '9|9']`\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|6954\t | abdff2be-bba9-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|12535\t| 0748041e-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|12691\t| 5aa7100a-bbb6-11e8-b2ba-ac1f6b6435d0\t|9\\|9 |[9] |\n|...| ...\t|... |... |\n\nFilter images with 1 label:\n\n`df[df['Label_lst'].map(len) == 1]`\n\n||ID |Label|Label_lst|\n| --- | --- |--- |\n|5| 5e22a522-bb99-11e8-b2b9-ac1f6b6435d0\t|0 |[0] |\n|6| 5f79a114-bb99-11e8-b2b9-ac1f6b6435d0\t|14 |[14] |\n|9| 5c801c04-bb99-11e8-b2b9-ac1f6b6435d0\t|14 |[14] |\n|...| ...\t|... |... |\n",
    "1211892": "sorry, I missed a bracket in my explanation. The problem I'm having still holds. test = df[~df['Label'].str.contains('|')] just gives me an empty dataframe...",
    "1212133": ""
  }
}