{
  "id": 166660,
  "title": "Combatting class imbalance question",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166660",
  "author_name": "Caleb Woy",
  "post_date": "2020-07-13T15:36:44.749000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm trying to make a weighted random sampler for one of my CNNs. To do so I need to loop through the dataset and get a list of all the labels that I can then replace with their respective weights. I'm currently doing it like this:</p>\n\n<p>class_arr = [0 if label == 0 else 1 for _, label in train_set]</p>\n\n<p>train_set is a random split of an ImageFolder that holds all my training images. It's very slow... anyone know a faster way to accomplish this?</p>",
  "messages": [
    {
      "id": 928364,
      "postDate": "2020-07-13T23:24:30.700Z",
      "content": "<p>I guess if you switch to numpy arrays, it can be faster.</p>\n\n<p>a = np.array([0.0,0.1,0.9,0.0,0.4])\nprint(np.where(a&gt;0,1,0))</p>",
      "rawMarkdown": "I guess if you switch to numpy arrays, it can be faster.\n\na = np.array([0.0,0.1,0.9,0.0,0.4])\nprint(np.where(a&gt;0,1,0))",
      "replies": [
        {
          "id": 928448,
          "postDate": "2020-07-14T02:07:25.737Z",
          "content": "<p>good idea. I think I'll try that next time.</p>",
          "rawMarkdown": "good idea. I think I'll try that next time."
        }
      ]
    },
    {
      "id": 927963,
      "postDate": "2020-07-13T16:54:51.063Z",
      "content": "<p>If you're using TensorFlow/Keras, you can just add weights to the fit call</p>\n\n<pre><code>model.fit(X,y,class_weight = {0:1,1:2})\n</code></pre>",
      "rawMarkdown": "If you're using TensorFlow/Keras, you can just add weights to the fit call\n\n    model.fit(X,y,class_weight = {0:1,1:2})",
      "replies": [
        {
          "id": 928039,
          "postDate": "2020-07-13T17:50:44.593Z",
          "content": "<p>Thanks for the response Chris, however, I'm using pytorch.</p>",
          "rawMarkdown": "Thanks for the response Chris, however, I'm using pytorch."
        },
        {
          "id": 928049,
          "postDate": "2020-07-13T18:02:00.253Z",
          "content": "<p>There's probably an easy way but i don't know it.</p>",
          "rawMarkdown": "There's probably an easy way but i don't know it."
        },
        {
          "id": 928090,
          "postDate": "2020-07-13T18:44:07.687Z",
          "content": "<p>That's okay. Patience is a virtue.</p>",
          "rawMarkdown": "That's okay. Patience is a virtue."
        }
      ]
    },
    {
      "id": 927806,
      "postDate": "2020-07-13T15:36:44.750Z",
      "content": "<p>I'm trying to make a weighted random sampler for one of my CNNs. To do so I need to loop through the dataset and get a list of all the labels that I can then replace with their respective weights. I'm currently doing it like this:</p>\n\n<p>class_arr = [0 if label == 0 else 1 for _, label in train_set]</p>\n\n<p>train_set is a random split of an ImageFolder that holds all my training images. It's very slow... anyone know a faster way to accomplish this?</p>",
      "rawMarkdown": "I'm trying to make a weighted random sampler for one of my CNNs. To do so I need to loop through the dataset and get a list of all the labels that I can then replace with their respective weights. I'm currently doing it like this:\n\nclass_arr = [0 if label == 0 else 1 for _, label in train_set]\n\ntrain_set is a random split of an ImageFolder that holds all my training images. It's very slow... anyone know a faster way to accomplish this?"
    },
    {
      "id": 928459,
      "postDate": "2020-07-14T02:46:19.503Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 928554,
          "postDate": "2020-07-14T04:43:58.927Z",
          "content": "<p>I need to get this list of labels to make a WeightedRandomSampler to pass to a BatchSampler to pass to a DataLoader. The dataloader class just holds a list of indices in the dataset it is an iterable for. I don't think creating an extra dataloader will make this any faster. </p>",
          "rawMarkdown": "I need to get this list of labels to make a WeightedRandomSampler to pass to a BatchSampler to pass to a DataLoader. The dataloader class just holds a list of indices in the dataset it is an iterable for. I don't think creating an extra dataloader will make this any faster. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 928364,
      "author_name": "ELEVEN",
      "author_url": "",
      "post_date": "2020-07-13T23:24:30.700000",
      "content": "<p>I guess if you switch to numpy arrays, it can be faster.</p>\n\n<p>a = np.array([0.0,0.1,0.9,0.0,0.4])\nprint(np.where(a&gt;0,1,0))</p>",
      "votes": 0,
      "replies": [
        {
          "id": 928448,
          "author_name": "Caleb Woy",
          "author_url": "",
          "post_date": "2020-07-14T02:07:25.737000",
          "content": "<p>good idea. I think I'll try that next time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 927963,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-13T16:54:51.063000",
      "content": "<p>If you're using TensorFlow/Keras, you can just add weights to the fit call</p>\n\n<pre><code>model.fit(X,y,class_weight = {0:1,1:2})\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 928039,
          "author_name": "Caleb Woy",
          "author_url": "",
          "post_date": "2020-07-13T17:50:44.593000",
          "content": "<p>Thanks for the response Chris, however, I'm using pytorch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 928049,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-13T18:02:00.253000",
          "content": "<p>There's probably an easy way but i don't know it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 928090,
          "author_name": "Caleb Woy",
          "author_url": "",
          "post_date": "2020-07-13T18:44:07.687000",
          "content": "<p>That's okay. Patience is a virtue.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 928459,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T02:46:19.503000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 928554,
          "author_name": "Caleb Woy",
          "author_url": "",
          "post_date": "2020-07-14T04:43:58.927000",
          "content": "<p>I need to get this list of labels to make a WeightedRandomSampler to pass to a BatchSampler to pass to a DataLoader. The dataloader class just holds a list of indices in the dataset it is an iterable for. I don't think creating an extra dataloader will make this any faster. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "928364": "I guess if you switch to numpy arrays, it can be faster.\n\na = np.array([0.0,0.1,0.9,0.0,0.4])\nprint(np.where(a&gt;0,1,0))",
    "927963": "If you're using TensorFlow/Keras, you can just add weights to the fit call\n\n    model.fit(X,y,class_weight = {0:1,1:2})",
    "927806": "I'm trying to make a weighted random sampler for one of my CNNs. To do so I need to loop through the dataset and get a list of all the labels that I can then replace with their respective weights. I'm currently doing it like this:\n\nclass_arr = [0 if label == 0 else 1 for _, label in train_set]\n\ntrain_set is a random split of an ImageFolder that holds all my training images. It's very slow... anyone know a faster way to accomplish this?",
    "928459": ""
  }
}