{
  "id": 112245,
  "title": "Why my dices are the same when I optimized thresholds ",
  "url": "/competitions/understanding_cloud_organization/discussion/112245",
  "author_name": "",
  "post_date": "2019-10-11T16:23:12.385140500Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>```\nbatch_size = 4\nprobabilities = np.zeros((len(train_ids)*4, 350, 525), dtype = np.float32)\ntruth_mask = np.zeros((len(train_ids)*batch_size, 350, 525), dtype = np.float32)</p>\n\n<p>total_batches = len(dataloader)\ntk0 = tqdm(dataloader, total=total_batches)\nwith torch.no_grad():\n    try:\n        with tk0 as t:\n            for itr, batch in enumerate(t):\n                nnnn, images, targets, _ = batch #targets:2,4,525,350\n                mean = torch.tensor([0.485, 0.456, 0.406])[:,None,None].cuda()\n                std = torch.tensor([0.229, 0.224, 0.225])[:,None,None].cuda()\n                images = (images.cuda() - mean)/std  #2,3,525,350\n                outputs = model(images) #2,4,525,350\n                outputs = torch.sigmoid(outputs).cpu().detach().numpy()\n                probabilities[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = outputs.reshape(-1,350,525)\n                truth_mask[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = targets.numpy().reshape(-1,350,525)\n    except KeyboardInterrupt:\n        t.close()\n        raise\n    t.close()</p>\n\n<p>class_params = {}\nfor class_id in range(4):\n    print(class_id)\n    attempts = []\n    for t in range(0, 100, 5):\n        t /= 100\n        for ms in [2000, 5000, 10000, 15000, 20000, 22500, 25000]:\n            masks = []\n            for i in range(class_id, len(probabilities), 4): #(4444, 350, 525) \n                #probabilities[i]:(350,525)\n                predict, num_predict = post_process(probabilities[i], t, ms)\n                probabilities[i] = predict</p>\n\n<pre><code>        d = []\n        for ii, jj in zip(probabilities[class_id::4], truth_mask[class_id::4]):\n            #pdb.set_trace()\n           if (ii.sum() == 0) &amp; (jj.sum() == 0):\n                d.append(1)\n           else:\n                d.append(dice_valid(ii, jj))\n\n        attempts.append((t, ms, np.mean(d)))\n    print(ms, 'finished')\n\nattempts_df = pd.DataFrame(attempts, columns=['threshold', 'size', 'dice'])\n\n\nattempts_df = attempts_df.sort_values('dice', ascending=False)\nprint(attempts_df.head())\nbest_threshold = attempts_df['threshold'].values[0]\nbest_size = attempts_df['size'].values[0]\n\nclass_params[class_id] = (best_threshold, best_size)\n</code></pre>\n\n<p>```\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F675401%2F24daeec0082652c49eb32b0c903ab0b1%2F36020191012002108069.jpg?generation=1570810977255932&amp;alt=media\" alt=\"\"></p>\n\n<p>This code output the same dice with every class type.Why my dices are the same when I optimized thresholds?</p>",
  "messages": [
    {
      "id": "646724",
      "postDate": "10/11/2019 16:23:12",
      "content": "<p>```\nbatch_size = 4\nprobabilities = np.zeros((len(train_ids)*4, 350, 525), dtype = np.float32)\ntruth_mask = np.zeros((len(train_ids)*batch_size, 350, 525), dtype = np.float32)</p>\n\n<p>total_batches = len(dataloader)\ntk0 = tqdm(dataloader, total=total_batches)\nwith torch.no_grad():\n    try:\n        with tk0 as t:\n            for itr, batch in enumerate(t):\n                nnnn, images, targets, _ = batch #targets:2,4,525,350\n                mean = torch.tensor([0.485, 0.456, 0.406])[:,None,None].cuda()\n                std = torch.tensor([0.229, 0.224, 0.225])[:,None,None].cuda()\n                images = (images.cuda() - mean)/std  #2,3,525,350\n                outputs = model(images) #2,4,525,350\n                outputs = torch.sigmoid(outputs).cpu().detach().numpy()\n                probabilities[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = outputs.reshape(-1,350,525)\n                truth_mask[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = targets.numpy().reshape(-1,350,525)\n    except KeyboardInterrupt:\n        t.close()\n        raise\n    t.close()</p>\n\n<p>class_params = {}\nfor class_id in range(4):\n    print(class_id)\n    attempts = []\n    for t in range(0, 100, 5):\n        t /= 100\n        for ms in [2000, 5000, 10000, 15000, 20000, 22500, 25000]:\n            masks = []\n            for i in range(class_id, len(probabilities), 4): #(4444, 350, 525) \n                #probabilities[i]:(350,525)\n                predict, num_predict = post_process(probabilities[i], t, ms)\n                probabilities[i] = predict</p>\n\n<pre><code>        d = []\n        for ii, jj in zip(probabilities[class_id::4], truth_mask[class_id::4]):\n            #pdb.set_trace()\n           if (ii.sum() == 0) &amp; (jj.sum() == 0):\n                d.append(1)\n           else:\n                d.append(dice_valid(ii, jj))\n\n        attempts.append((t, ms, np.mean(d)))\n    print(ms, 'finished')\n\nattempts_df = pd.DataFrame(attempts, columns=['threshold', 'size', 'dice'])\n\n\nattempts_df = attempts_df.sort_values('dice', ascending=False)\nprint(attempts_df.head())\nbest_threshold = attempts_df['threshold'].values[0]\nbest_size = attempts_df['size'].values[0]\n\nclass_params[class_id] = (best_threshold, best_size)\n</code></pre>\n\n<p>```\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F675401%2F24daeec0082652c49eb32b0c903ab0b1%2F36020191012002108069.jpg?generation=1570810977255932&amp;alt=media\" alt=\"\"></p>\n\n<p>This code output the same dice with every class type.Why my dices are the same when I optimized thresholds?</p>",
      "rawMarkdown": "```\nbatch_size = 4\nprobabilities = np.zeros((len(train_ids)*4, 350, 525), dtype = np.float32)\ntruth_mask = np.zeros((len(train_ids)*batch_size, 350, 525), dtype = np.float32)\n\ntotal_batches = len(dataloader)\ntk0 = tqdm(dataloader, total=total_batches)\nwith torch.no_grad():\n    try:\n        with tk0 as t:\n            for itr, batch in enumerate(t):\n                nnnn, images, targets, _ = batch #targets:2,4,525,350\n                mean = torch.tensor([0.485, 0.456, 0.406])[:,None,None].cuda()\n                std = torch.tensor([0.229, 0.224, 0.225])[:,None,None].cuda()\n                images = (images.cuda() - mean)/std  #2,3,525,350\n                outputs = model(images) #2,4,525,350\n                outputs = torch.sigmoid(outputs).cpu().detach().numpy()\n                probabilities[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = outputs.reshape(-1,350,525)\n                truth_mask[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = targets.numpy().reshape(-1,350,525)\n    except KeyboardInterrupt:\n        t.close()\n        raise\n    t.close()\n\nclass_params = {}\nfor class_id in range(4):\n    print(class_id)\n    attempts = []\n    for t in range(0, 100, 5):\n        t /= 100\n        for ms in [2000, 5000, 10000, 15000, 20000, 22500, 25000]:\n            masks = []\n            for i in range(class_id, len(probabilities), 4): #(4444, 350, 525) \n                #probabilities[i]:(350,525)\n                predict, num_predict = post_process(probabilities[i], t, ms)\n                probabilities[i] = predict\n            \n            d = []\n            for ii, jj in zip(probabilities[class_id::4], truth_mask[class_id::4]):\n                #pdb.set_trace()\n               if (ii.sum() == 0) &amp; (jj.sum() == 0):\n                    d.append(1)\n               else:\n                    d.append(dice_valid(ii, jj))\n\n            attempts.append((t, ms, np.mean(d)))\n        print(ms, 'finished')\n\n    attempts_df = pd.DataFrame(attempts, columns=['threshold', 'size', 'dice'])\n\n\n    attempts_df = attempts_df.sort_values('dice', ascending=False)\n    print(attempts_df.head())\n    best_threshold = attempts_df['threshold'].values[0]\n    best_size = attempts_df['size'].values[0]\n    \n    class_params[class_id] = (best_threshold, best_size)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F675401%2F24daeec0082652c49eb32b0c903ab0b1%2F36020191012002108069.jpg?generation=1570810977255932&amp;alt=media)\n\nThis code output the same dice with every class type.Why my dices are the same when I optimized thresholds?",
      "votes": null
    },
    {
      "id": "647095",
      "postDate": "10/12/2019 04:09:29",
      "content": "<p>I may recommend you check truth mask and prediction tensor shape. </p>",
      "rawMarkdown": "I may recommend you check truth mask and prediction tensor shape.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 647095,
      "author_name": "miklgr500",
      "author_url": "",
      "post_date": "10/12/2019 04:09:29",
      "content": "<p>I may recommend you check truth mask and prediction tensor shape. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "646724": "```\nbatch_size = 4\nprobabilities = np.zeros((len(train_ids)*4, 350, 525), dtype = np.float32)\ntruth_mask = np.zeros((len(train_ids)*batch_size, 350, 525), dtype = np.float32)\n\ntotal_batches = len(dataloader)\ntk0 = tqdm(dataloader, total=total_batches)\nwith torch.no_grad():\n    try:\n        with tk0 as t:\n            for itr, batch in enumerate(t):\n                nnnn, images, targets, _ = batch #targets:2,4,525,350\n                mean = torch.tensor([0.485, 0.456, 0.406])[:,None,None].cuda()\n                std = torch.tensor([0.229, 0.224, 0.225])[:,None,None].cuda()\n                images = (images.cuda() - mean)/std  #2,3,525,350\n                outputs = model(images) #2,4,525,350\n                outputs = torch.sigmoid(outputs).cpu().detach().numpy()\n                probabilities[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = outputs.reshape(-1,350,525)\n                truth_mask[itr*batch_size*4:(itr+1)*batch_size*4, :, :] = targets.numpy().reshape(-1,350,525)\n    except KeyboardInterrupt:\n        t.close()\n        raise\n    t.close()\n\nclass_params = {}\nfor class_id in range(4):\n    print(class_id)\n    attempts = []\n    for t in range(0, 100, 5):\n        t /= 100\n        for ms in [2000, 5000, 10000, 15000, 20000, 22500, 25000]:\n            masks = []\n            for i in range(class_id, len(probabilities), 4): #(4444, 350, 525) \n                #probabilities[i]:(350,525)\n                predict, num_predict = post_process(probabilities[i], t, ms)\n                probabilities[i] = predict\n            \n            d = []\n            for ii, jj in zip(probabilities[class_id::4], truth_mask[class_id::4]):\n                #pdb.set_trace()\n               if (ii.sum() == 0) &amp; (jj.sum() == 0):\n                    d.append(1)\n               else:\n                    d.append(dice_valid(ii, jj))\n\n            attempts.append((t, ms, np.mean(d)))\n        print(ms, 'finished')\n\n    attempts_df = pd.DataFrame(attempts, columns=['threshold', 'size', 'dice'])\n\n\n    attempts_df = attempts_df.sort_values('dice', ascending=False)\n    print(attempts_df.head())\n    best_threshold = attempts_df['threshold'].values[0]\n    best_size = attempts_df['size'].values[0]\n    \n    class_params[class_id] = (best_threshold, best_size)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F675401%2F24daeec0082652c49eb32b0c903ab0b1%2F36020191012002108069.jpg?generation=1570810977255932&amp;alt=media)\n\nThis code output the same dice with every class type.Why my dices are the same when I optimized thresholds?",
    "647095": "I may recommend you check truth mask and prediction tensor shape."
  },
  "source": "meta"
}