{
  "id": 73938,
  "title": "How I found myself in the \"gold\" zone",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73938",
  "author_name": "Thundo",
  "post_date": "2018-12-06T20:56:15.053000",
  "votes": 56,
  "comment_count": 108,
  "views": 0,
  "content": "<p>After last rescoring and a couple more submissions I found myself in the top 10. It's not my first competition, but I really wonder how I got so high, even if only temporarily and probably overfitting the public leaderbord. </p>\n\n<p>I want to share what are my best \"selling points\" right now, hoping to get some feedback from others:\n- strong experiment pipeline\n- off-the-shelf pretrained model\n- \"standard\" augmentation\n- HPA external data\n- cross-validation ensemble\n- multilabel stratification\n- weighted samplers\n- lr schedule\n- image size 512\n- \"flat\" threshold\n- no tta</p>\n\n<p>As I said before, I may be overfitting the public LB badly, so take everything with a pinch of salt....\nThank you!</p>",
  "messages": [
    {
      "id": 434718,
      "postDate": "2018-12-06T20:56:15.053Z",
      "content": "<p>After last rescoring and a couple more submissions I found myself in the top 10. It's not my first competition, but I really wonder how I got so high, even if only temporarily and probably overfitting the public leaderbord. </p>\n\n<p>I want to share what are my best \"selling points\" right now, hoping to get some feedback from others:\n- strong experiment pipeline\n- off-the-shelf pretrained model\n- \"standard\" augmentation\n- HPA external data\n- cross-validation ensemble\n- multilabel stratification\n- weighted samplers\n- lr schedule\n- image size 512\n- \"flat\" threshold\n- no tta</p>\n\n<p>As I said before, I may be overfitting the public LB badly, so take everything with a pinch of salt....\nThank you!</p>",
      "rawMarkdown": "After last rescoring and a couple more submissions I found myself in the top 10. It's not my first competition, but I really wonder how I got so high, even if only temporarily and probably overfitting the public leaderbord. \n\nI want to share what are my best \"selling points\" right now, hoping to get some feedback from others:\n- strong experiment pipeline\n- off-the-shelf pretrained model\n- \"standard\" augmentation\n- HPA external data\n- cross-validation ensemble\n- multilabel stratification\n- weighted samplers\n- lr schedule\n- image size 512\n- \"flat\" threshold\n- no tta\n\nAs I said before, I may be overfitting the public LB badly, so take everything with a pinch of salt....\nThank you!",
      "votes": 57
    },
    {
      "id": 435859,
      "postDate": "2018-12-08T23:11:30.827Z",
      "content": "<p>Hey Thundo I found myself in a similar spot! </p>\n\n<ul>\n<li><p>The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.</p></li>\n<li><p>Image size 512 seems to be enough to get a decent score</p></li>\n<li><p>If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)</p></li>\n<li><p>I'm unsure about how the distribution of classes looks right now, as the correction of the leak moved rare cases from private to public (possible overfitting to public leaderboard as there are less rare cases in the private test set)</p></li>\n</ul>\n\n<p>Saying all this I am very much new to deep learning so please correct me if I make wrong assumptions</p>",
      "rawMarkdown": "Hey Thundo I found myself in a similar spot! \n\n* The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.\n\n* Image size 512 seems to be enough to get a decent score\n\n* If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)\n\n* I'm unsure about how the distribution of classes looks right now, as the correction of the leak moved rare cases from private to public (possible overfitting to public leaderboard as there are less rare cases in the private test set)\n\nSaying all this I am very much new to deep learning so please correct me if I make wrong assumptions",
      "votes": 5,
      "replies": [
        {
          "id": 435868,
          "postDate": "2018-12-08T23:51:25.353Z",
          "content": "<p>Hi I was wondering loss functions are you using, and how are you balancing your dataset? I'm using Focal + F1, and balancing by giving each sample a score of 1/proportion_of_class for each positive class in the sample then taking the mean. Though my validation loss is always a lot higher than my training loss, even w/ image augmentation.</p>",
          "rawMarkdown": "Hi I was wondering loss functions are you using, and how are you balancing your dataset? I'm using Focal + F1, and balancing by giving each sample a score of 1/proportion_of_class for each positive class in the sample then taking the mean. Though my validation loss is always a lot higher than my training loss, even w/ image augmentation."
        },
        {
          "id": 435884,
          "postDate": "2018-12-09T01:31:11.550Z",
          "content": "<p>I'm using Focalloss and do balancing mostly with thresholding my output (for each class) </p>\n\n<p>Also using a validation split that has a similar class imbalance helps I think</p>",
          "rawMarkdown": "I'm using Focalloss and do balancing mostly with thresholding my output (for each class) \n\nAlso using a validation split that has a similar class imbalance helps I think"
        },
        {
          "id": 435955,
          "postDate": "2018-12-09T06:28:07.347Z",
          "content": "<p>I see so you don't explicitly do any under/oversampling of each sample?</p>",
          "rawMarkdown": "I see so you don't explicitly do any under/oversampling of each sample?"
        },
        {
          "id": 436000,
          "postDate": "2018-12-09T08:53:49.907Z",
          "content": "<p>&gt; The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.</p>\n\n<p>You're right. Considering download difficulties, increased storage (original and resized versions) and training time, it could be easily an entry barrier for lots of users.</p>\n\n<p>&gt; Image size 512 seems to be enough to get a decent score</p>\n\n<p>I would like to try something bigger but I think I don't have enough resources at my disposal.</p>\n\n<p>&gt; If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)</p>\n\n<p>I tried thresholds based on the predictions of the out-of-fold validation set. The CV score is better but not the LB. Maybe as someone suggested, this is a byproduct of the weighted sampling or maybe just overfitting.</p>",
          "rawMarkdown": "&gt; The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.\n\nYou're right. Considering download difficulties, increased storage (original and resized versions) and training time, it could be easily an entry barrier for lots of users.\n\n&gt; Image size 512 seems to be enough to get a decent score\n\nI would like to try something bigger but I think I don't have enough resources at my disposal.\n\n&gt; If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)\n\nI tried thresholds based on the predictions of the out-of-fold validation set. The CV score is better but not the LB. Maybe as someone suggested, this is a byproduct of the weighted sampling or maybe just overfitting.\n",
          "votes": 1
        },
        {
          "id": 436539,
          "postDate": "2018-12-10T13:52:56.127Z",
          "content": "<p>Nice share! \nHmm, Thundo, could you please tell me where I can get those \"external\" data ? \nI haven't read any info about that on the \"Data\" column of the competition.</p>",
          "rawMarkdown": "Nice share! \nHmm, Thundo, could you please tell me where I can get those \"external\" data ? \nI haven't read any info about that on the \"Data\" column of the competition.",
          "votes": 1
        },
        {
          "id": 436547,
          "postDate": "2018-12-10T14:06:01.227Z",
          "content": "<p>You can find instructions to download the external data here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#432870\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#432870</a></p>",
          "rawMarkdown": "You can find instructions to download the external data here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#432870",
          "votes": 1
        }
      ]
    },
    {
      "id": 434983,
      "postDate": "2018-12-07T09:05:23.300Z",
      "content": "<p>I use TTA (standard), independent thresholds, no-ensemble (at the moment), v1 HPA external data. I think finding the right thresholds is important and for sure would help you. No easy task though.</p>",
      "rawMarkdown": "I use TTA (standard), independent thresholds, no-ensemble (at the moment), v1 HPA external data. I think finding the right thresholds is important and for sure would help you. No easy task though.",
      "votes": 6,
      "replies": [
        {
          "id": 435076,
          "postDate": "2018-12-07T13:10:24.403Z",
          "content": "<p>Thanks <a href=\"/arnaurm\">@arnaurm</a>. </p>\n\n<p>I plan to add TTA next. Could you share more about your take on threshold selection?</p>",
          "rawMarkdown": "Thanks @arnaurm. \n\nI plan to add TTA next. Could you share more about your take on threshold selection?"
        },
        {
          "id": 437696,
          "postDate": "2018-12-12T10:20:39.800Z",
          "content": "<p>I just find the thresholds that fit best with my validation data. Later maybe I will try an ensemble with majority voting. Having said this, don't trust the public leaderboard, I have the feeling, as you said, that some of us are overfitting with the leaked data. I think real score is inbetween 0.59-0.60. </p>",
          "rawMarkdown": "I just find the thresholds that fit best with my validation data. Later maybe I will try an ensemble with majority voting. Having said this, don't trust the public leaderboard, I have the feeling, as you said, that some of us are overfitting with the leaked data. I think real score is inbetween 0.59-0.60. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 436979,
      "postDate": "2018-12-11T07:11:16.363Z",
      "content": "<p>Thanks for your share. I have a question:\nHow many channels do you used, I found I got a higher score with 3 channels(RGB) than with 4 channels(RGBY).</p>",
      "rawMarkdown": "Thanks for your share. I have a question:\nHow many channels do you used, I found I got a higher score with 3 channels(RGB) than with 4 channels(RGBY).",
      "votes": 3,
      "replies": [
        {
          "id": 437658,
          "postDate": "2018-12-12T09:35:49.547Z",
          "content": "<p>I tried both approaches a while back... In the end I resorted to RGBY.</p>",
          "rawMarkdown": "I tried both approaches a while back... In the end I resorted to RGBY.",
          "votes": 1
        },
        {
          "id": 437690,
          "postDate": "2018-12-12T10:12:27.290Z",
          "content": "<p>The challenge is the the HPA images only has 3 layers. Do you add a yellow layer filled with 0. or null vlaues to make the shape consistent with the competition data images?</p>",
          "rawMarkdown": "The challenge is the the HPA images only has 3 layers. Do you add a yellow layer filled with 0. or null vlaues to make the shape consistent with the competition data images?"
        },
        {
          "id": 437707,
          "postDate": "2018-12-12T10:45:07.673Z",
          "content": "<p>That's what I thought at the beginning... Actually most of them have 4 channels... You can find more explanations and even working code in <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">the external data thread</a></p>",
          "rawMarkdown": "That's what I thought at the beginning... Actually most of them have 4 channels... You can find more explanations and even working code in [the external data thread](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984)",
          "votes": 1
        }
      ]
    },
    {
      "id": 440769,
      "postDate": "2018-12-18T01:04:18.513Z",
      "content": "<p>I'm just wondering if anyone has successfully gone over 512 in image size. It looks do-able to me if the batch size is not large. Of course, it'll be very slow...</p>",
      "rawMarkdown": "I'm just wondering if anyone has successfully gone over 512 in image size. It looks do-able to me if the batch size is not large. Of course, it'll be very slow...",
      "votes": 1,
      "replies": [
        {
          "id": 440783,
          "postDate": "2018-12-18T01:17:33.593Z",
          "content": "<p>I attempted 1024, but was not able to run enough epochs to make it competitive with my 512 solution. The constraint really is time, and I think most people (who are operating with constrained resources) would be better served continuing to experiment at 512 and building ensembles rather than trying the higher resolution.</p>",
          "rawMarkdown": "I attempted 1024, but was not able to run enough epochs to make it competitive with my 512 solution. The constraint really is time, and I think most people (who are operating with constrained resources) would be better served continuing to experiment at 512 and building ensembles rather than trying the higher resolution."
        },
        {
          "id": 441641,
          "postDate": "2018-12-18T22:02:51.457Z",
          "content": "<p>I tried something in-between 512 and 1024. Training is painfully slow on my GTX 1070 and it didn't improve score...</p>",
          "rawMarkdown": "I tried something in-between 512 and 1024. Training is painfully slow on my GTX 1070 and it didn't improve score..."
        },
        {
          "id": 446525,
          "postDate": "2018-12-28T08:36:24.930Z",
          "content": "<p>Batch size is also too small and it actually affects convergence. \nI think that one needs really expensive GPUs to train deep CNNs on 1024x1024 pictures with a decent batch size.</p>",
          "rawMarkdown": "Batch size is also too small and it actually affects convergence. \nI think that one needs really expensive GPUs to train deep CNNs on 1024x1024 pictures with a decent batch size."
        }
      ]
    },
    {
      "id": 435443,
      "postDate": "2018-12-08T03:35:10.507Z",
      "content": "<p>What do you mean by \"cross-validation ensemble\", can you help with some example? Thanks!</p>",
      "rawMarkdown": "What do you mean by \"cross-validation ensemble\", can you help with some example? Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 435997,
          "postDate": "2018-12-09T08:43:53.713Z",
          "content": "<p>I use a 5-fold CV for training, then I ensemble the predictions of the 5 different models.</p>",
          "rawMarkdown": "I use a 5-fold CV for training, then I ensemble the predictions of the 5 different models."
        },
        {
          "id": 436076,
          "postDate": "2018-12-09T14:27:35.033Z",
          "content": "<p>Thank You!</p>",
          "rawMarkdown": "Thank You!"
        },
        {
          "id": 436378,
          "postDate": "2018-12-10T07:54:32.830Z",
          "content": "<p>hi, how do you ensemble predictions from 5-fold cv models? Do you average all 5 probability on test set? If so, what thresholds do you use? Since 5-fold have differnet best threshold on validation set, it is hard to choose the best threshold on test set.</p>",
          "rawMarkdown": "hi, how do you ensemble predictions from 5-fold cv models? Do you average all 5 probability on test set? If so, what thresholds do you use? Since 5-fold have differnet best threshold on validation set, it is hard to choose the best threshold on test set."
        },
        {
          "id": 436635,
          "postDate": "2018-12-10T17:05:25.200Z",
          "content": "<p>Yes, I use the mean of the predictions. \nThresholds on the out of fold are not so different for me, so I averaged them too.\nBut as I wrote I use a flat threshold for every class.</p>",
          "rawMarkdown": "Yes, I use the mean of the predictions. \nThresholds on the out of fold are not so different for me, so I averaged them too.\nBut as I wrote I use a flat threshold for every class."
        }
      ]
    },
    {
      "id": 434834,
      "postDate": "2018-12-07T02:44:43.477Z",
      "content": "<p>Thanks for the post and congratulations.\nI use a tilted threshold. Perhaps the flat works for you due to the weighted sampler.</p>\n\n<p>Are you modifying your submission according to the HPA leak? Also, are you training on the images? I am trying that now.</p>",
      "rawMarkdown": "Thanks for the post and congratulations.\nI use a tilted threshold. Perhaps the flat works for you due to the weighted sampler.\n\nAre you modifying your submission according to the HPA leak? Also, are you training on the images? I am trying that now.\n",
      "votes": 1,
      "replies": [
        {
          "id": 434959,
          "postDate": "2018-12-07T08:10:21.820Z",
          "content": "<p>I tried per-class thresholds using the validation set but got worse results. This could, of course, entirely be related to the public leaderboard distribution.</p>\n\n<p>I use HPA as an extension of my train dataset, so yes, I use it for training. It's hard to say, due the rescoring, how much improved my LB score.</p>",
          "rawMarkdown": "I tried per-class thresholds using the validation set but got worse results. This could, of course, entirely be related to the public leaderboard distribution.\n\nI use HPA as an extension of my train dataset, so yes, I use it for training. It's hard to say, due the rescoring, how much improved my LB score."
        },
        {
          "id": 435418,
          "postDate": "2018-12-08T02:28:41.310Z",
          "content": "<p>hello! pete,do you mean that you can get so high LB score when you don't use  v1 HPA external data.</p>",
          "rawMarkdown": "hello! pete,do you mean that you can get so high LB score when you don't use  v1 HPA external data."
        },
        {
          "id": 435876,
          "postDate": "2018-12-09T00:10:17.107Z",
          "content": "<p>I had .504 before the adjustment without HPS usage. After the adjustment, I had .503 (reasonable). Then, I used the HPA data to take advantage of the leak and got 0.55 or so. That was ONLY using the leak csv file posted - none of the images - to fix up the submission file. So far, I have not used the HPA images directly though I plan to do it for training augmentation.</p>",
          "rawMarkdown": "I had .504 before the adjustment without HPS usage. After the adjustment, I had .503 (reasonable). Then, I used the HPA data to take advantage of the leak and got 0.55 or so. That was ONLY using the leak csv file posted - none of the images - to fix up the submission file. So far, I have not used the HPA images directly though I plan to do it for training augmentation."
        },
        {
          "id": 436327,
          "postDate": "2018-12-10T06:07:43.610Z",
          "content": "<p>you just used the leak csv file posted to fix up the submission file? is it permitted?</p>",
          "rawMarkdown": "you just used the leak csv file posted to fix up the submission file? is it permitted?"
        },
        {
          "id": 436383,
          "postDate": "2018-12-10T08:06:47.153Z",
          "content": "<p>Yes, just that file. It is permitted. It does not give anyone an advantage because it is public. And it does not appear to detract from the challenge goals. It is worth thinking about what happens on private leaderboard - I have not fully understood the organizer comments on the rare classes.</p>",
          "rawMarkdown": "Yes, just that file. It is permitted. It does not give anyone an advantage because it is public. And it does not appear to detract from the challenge goals. It is worth thinking about what happens on private leaderboard - I have not fully understood the organizer comments on the rare classes."
        },
        {
          "id": 436414,
          "postDate": "2018-12-10T08:58:00.730Z",
          "content": "<p>My understanding is that anyone who can't replicate their public LB score with the leak file by a submission that excludes it is in for a fall on the private LB.</p>",
          "rawMarkdown": "My understanding is that anyone who can't replicate their public LB score with the leak file by a submission that excludes it is in for a fall on the private LB.",
          "votes": 2
        },
        {
          "id": 437816,
          "postDate": "2018-12-12T14:53:05.887Z",
          "content": "<p>oh,very good , thank you for your useful answer.can you tell me where is the csv file?</p>",
          "rawMarkdown": "oh,very good , thank you for your useful answer.can you tell me where is the csv file?"
        }
      ]
    },
    {
      "id": 434761,
      "postDate": "2018-12-06T22:39:26.627Z",
      "content": "<p>Calibrating Probability with Undersampling for Unbalanced Classification - <a href=\"https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf\">https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf</a></p>\n\n<p>I'm a noob on this matter but intuitively I think if you oversample (or undersample) your minority (majority) classes, you change their priors. If you choose one sample at random from an unbalanced dataset, the probability of getting a sample from a seldom class is quite low. With weighted sampler, the probability of choosing a sample at random from the same class  increases, thus the need for probability calibration of the network's output when it is trained with an artificially balanced dataset. I might be saying nonsense though. Hopefully, an expert on this subject can comment more on that....\nAnyways, still can't get oversampling to work well in my model</p>",
      "rawMarkdown": "Calibrating Probability with Undersampling for Unbalanced Classification - https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf\n\nI'm a noob on this matter but intuitively I think if you oversample (or undersample) your minority (majority) classes, you change their priors. If you choose one sample at random from an unbalanced dataset, the probability of getting a sample from a seldom class is quite low. With weighted sampler, the probability of choosing a sample at random from the same class  increases, thus the need for probability calibration of the network's output when it is trained with an artificially balanced dataset. I might be saying nonsense though. Hopefully, an expert on this subject can comment more on that....\nAnyways, still can't get oversampling to work well in my model",
      "votes": 1,
      "replies": [
        {
          "id": 434966,
          "postDate": "2018-12-07T08:29:11.453Z",
          "content": "<p>Isn't the paper about exclusive classes? We have instead multi-labels. I think this change a bit the context, because the more popular labels don't \"disappear\" when you boost the low-impact ones, they may coexist. </p>\n\n<p>All in all I don't think the distribution changes so dramatically... Bottom line: I don't use any kind of postprocess on the probabilities.</p>",
          "rawMarkdown": "Isn't the paper about exclusive classes? We have instead multi-labels. I think this change a bit the context, because the more popular labels don't \"disappear\" when you boost the low-impact ones, they may coexist. \n\nAll in all I don't think the distribution changes so dramatically... Bottom line: I don't use any kind of postprocess on the probabilities."
        }
      ]
    },
    {
      "id": 434739,
      "postDate": "2018-12-06T21:49:10.043Z",
      "content": "<p>@Thundo Thank you for the tips! Do you calibrate your probabilities after using weighted samplers? I tried weighted sampler in my code but got no improvement in public LB (CV score was better though..not sure if I should trust it yet)</p>",
      "rawMarkdown": "@Thundo Thank you for the tips! Do you calibrate your probabilities after using weighted samplers? I tried weighted sampler in my code but got no improvement in public LB (CV score was better though..not sure if I should trust it yet)",
      "votes": 1,
      "replies": [
        {
          "id": 434746,
          "postDate": "2018-12-06T21:57:53.357Z",
          "content": "<p>What do you mean with \"calibrate your probabilities\"?</p>",
          "rawMarkdown": "What do you mean with \"calibrate your probabilities\"?"
        },
        {
          "id": 440107,
          "postDate": "2018-12-17T03:40:11.883Z",
          "content": "<p>I would suggest trying log dampening on the weights. You can find more info in this thread: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
          "rawMarkdown": "I would suggest trying log dampening on the weights. You can find more info in this thread: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065"
        }
      ]
    },
    {
      "id": 434736,
      "postDate": "2018-12-06T21:32:04.793Z",
      "content": "<p>Thanks Thundo for sharing those ideas... could you please explain what you call \"weighted samplers\" ? Do you have any kernel that describe this ? Thanks ! </p>",
      "rawMarkdown": "Thanks Thundo for sharing those ideas... could you please explain what you call \"weighted samplers\" ? Do you have any kernel that describe this ? Thanks ! ",
      "votes": 1,
      "replies": [
        {
          "id": 434738,
          "postDate": "2018-12-06T21:45:40.993Z",
          "content": "<p>Sorry, I don't have a kernel to show... I borrowed the term from PyTorch. I mean that I draw samples for the mini-batches taking class weigths into account...</p>",
          "rawMarkdown": "Sorry, I don't have a kernel to show... I borrowed the term from PyTorch. I mean that I draw samples for the mini-batches taking class weigths into account...",
          "votes": 1
        },
        {
          "id": 434809,
          "postDate": "2018-12-07T01:08:05.940Z",
          "content": "<p>Could you share 'weighted sampler' code in Pytorch for the multi-lables task? </p>",
          "rawMarkdown": "Could you share 'weighted sampler' code in Pytorch for the multi-lables task? "
        },
        {
          "id": 434963,
          "postDate": "2018-12-07T08:16:46.103Z",
          "content": "<p>PyTorch already has a <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\">weighted sampler</a>. What I do is simply set a weight for each training sample. </p>",
          "rawMarkdown": "PyTorch already has a [weighted sampler](https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler). What I do is simply set a weight for each training sample. ",
          "votes": 2
        },
        {
          "id": 436073,
          "postDate": "2018-12-09T14:11:44.893Z",
          "content": "<p>I was wondering what was the scheme that you used to decide what probability to sample each sample, due to the fact that each sample may have mulitple labels. I took the mean of the sums of the inverse counts for each label in the sample, but it doesn't seem to work very well...</p>",
          "rawMarkdown": "I was wondering what was the scheme that you used to decide what probability to sample each sample, due to the fact that each sample may have mulitple labels. I took the mean of the sums of the inverse counts for each label in the sample, but it doesn't seem to work very well..."
        },
        {
          "id": 440104,
          "postDate": "2018-12-17T03:36:13.743Z",
          "content": "<p>Do you find the best weights sample scheme, I think  per-class weights will not work well, because multi class coexists in one label, if use less class probility will increase large class weights too</p>",
          "rawMarkdown": "Do you find the best weights sample scheme, I think  per-class weights will not work well, because multi class coexists in one label, if use less class probility will increase large class weights too"
        },
        {
          "id": 443102,
          "postDate": "2018-12-21T02:45:17.117Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 439577,
      "postDate": "2018-12-15T20:42:57.920Z",
      "content": "<p>Nice share !</p>",
      "rawMarkdown": "Nice share !",
      "votes": 2
    },
    {
      "id": 439544,
      "postDate": "2018-12-15T18:59:51.957Z",
      "content": "<p>Using multilabel stratification, external data, and (log-dampened) weighted sampling got me from 0.473 -&gt; 0.494 public LB for my best single model (resnet50), no TTA.</p>",
      "rawMarkdown": "Using multilabel stratification, external data, and (log-dampened) weighted sampling got me from 0.473 -&gt; 0.494 public LB for my best single model (resnet50), no TTA.",
      "votes": 2,
      "replies": [
        {
          "id": 440111,
          "postDate": "2018-12-17T03:52:13.520Z",
          "content": "<p>Can you share how to learn this skill - multilabel stratification?</p>",
          "rawMarkdown": "Can you share how to learn this skill - multilabel stratification?"
        },
        {
          "id": 440119,
          "postDate": "2018-12-17T04:34:57.797Z",
          "content": "<p>You can one-hot encode the labels and call the method in  <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>. The usage is similar to cross validation split in ScikitLearn.  The pip/conda install might not work since the package is dated, just copy over the file (<a href=\"https://github.com/trent-b/iterative-stratification/blob/master/iterstrat/ml_stratifiers.py\">https://github.com/trent-b/iterative-stratification/blob/master/iterstrat/ml_stratifiers.py</a>) to your local workspace and it will work just fine.</p>",
          "rawMarkdown": "You can one-hot encode the labels and call the method in  https://github.com/trent-b/iterative-stratification. The usage is similar to cross validation split in ScikitLearn.  The pip/conda install might not work since the package is dated, just copy over the file (https://github.com/trent-b/iterative-stratification/blob/master/iterstrat/ml_stratifiers.py) to your local workspace and it will work just fine.\n",
          "votes": 2
        },
        {
          "id": 440121,
          "postDate": "2018-12-17T04:42:06.457Z",
          "content": "<p>The basic premise is that we want to do the train/validation split evenly across all labels. This is important in this competition since we have imbalance in dataset for different labels ( for ex. one end 12000 samples of one label(label=0) and on the other just 11(label=27)). The above package does just that.</p>",
          "rawMarkdown": "The basic premise is that we want to do the train/validation split evenly across all labels. This is important in this competition since we have imbalance in dataset for different labels ( for ex. one end 12000 samples of one label(label=0) and on the other just 11(label=27)). The above package does just that."
        },
        {
          "id": 440132,
          "postDate": "2018-12-17T05:35:19.587Z",
          "content": "<p>I get it, thank you~</p>",
          "rawMarkdown": "I get it, thank you~"
        },
        {
          "id": 441836,
          "postDate": "2018-12-19T06:32:29.193Z",
          "content": "<p>(log-dampened) weighted sampling？can you share it? i find weighted loss don't work well, the score of middle classes is lowest  </p>",
          "rawMarkdown": "(log-dampened) weighted sampling？can you share it? i find weighted loss don't work well, the score of middle classes is lowest  "
        },
        {
          "id": 442092,
          "postDate": "2018-12-19T13:28:33.107Z",
          "content": "<p>I used the code from here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
          "rawMarkdown": "I used the code from here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065"
        },
        {
          "id": 442452,
          "postDate": "2018-12-20T01:26:21.113Z",
          "content": "<p>i used the link without work, Maybe I have a problem with my weighting operation. i used pytorch with focal loss.you weight with'w' or 'pos_weight' parameter？ </p>",
          "rawMarkdown": "i used the link without work, Maybe I have a problem with my weighting operation. i used pytorch with focal loss.you weight with'w' or 'pos_weight' parameter？ "
        },
        {
          "id": 442520,
          "postDate": "2018-12-20T03:57:27.347Z",
          "content": "<p>I’ve used weighted sampling, but haven’t tried weighted loss yet. For weighted sampling I found the following thread very useful for Pytorch: <a href=\"https://discuss.pytorch.org/t/balanced-sampling-between-classes-with-torchvision-dataloader/2703/15\">https://discuss.pytorch.org/t/balanced-sampling-between-classes-with-torchvision-dataloader/2703/15</a></p>",
          "rawMarkdown": "I’ve used weighted sampling, but haven’t tried weighted loss yet. For weighted sampling I found the following thread very useful for Pytorch: https://discuss.pytorch.org/t/balanced-sampling-between-classes-with-torchvision-dataloader/2703/15"
        },
        {
          "id": 442621,
          "postDate": "2018-12-20T08:15:32.957Z",
          "content": "<p>thank you .but the link isn't existed.....</p>",
          "rawMarkdown": "thank you .but the link isn't existed....."
        },
        {
          "id": 442631,
          "postDate": "2018-12-20T08:40:07.433Z",
          "content": "<p><a href=\"/xu666bird\">@xu666bird</a> Remove the full-stop at the end of the link</p>",
          "rawMarkdown": "@xu666bird Remove the full-stop at the end of the link",
          "votes": 1
        },
        {
          "id": 442633,
          "postDate": "2018-12-20T08:47:31.377Z",
          "content": "<p>thank you . it's ok</p>",
          "rawMarkdown": "thank you . it's ok"
        },
        {
          "id": 442652,
          "postDate": "2018-12-20T09:25:53.987Z",
          "content": "<p><a href=\"/hortonhearsafoo\">@hortonhearsafoo</a> How did you convert your log-dampened weights to probabilities between 0 and 1 for your <code>WeightedRandomSampler</code>  because I see for you are not using log-dampened weights in your loss function and the log-dampened weights are in the range between 1.0 and 7.74 (As per the link you suggested to calculate weights above). My first assumption was you are using these weights in your loss function but I also got the worse results with class weights and Focal Loss (as pointed out by <a href=\"/xu666bird\">@xu666bird</a> ).  In WeightedRanomSampler we need to provide probabilities for each image. Could you please elaborate more on this? Thanks in advance :)</p>",
          "rawMarkdown": "@hortonhearsafoo How did you convert your log-dampened weights to probabilities between 0 and 1 for your `WeightedRandomSampler`  because I see for you are not using log-dampened weights in your loss function and the log-dampened weights are in the range between 1.0 and 7.74 (As per the link you suggested to calculate weights above). My first assumption was you are using these weights in your loss function but I also got the worse results with class weights and Focal Loss (as pointed out by @xu666bird ).  In WeightedRanomSampler we need to provide probabilities for each image. Could you please elaborate more on this? Thanks in advance :)"
        },
        {
          "id": 442821,
          "postDate": "2018-12-20T15:15:01.130Z",
          "content": "<p>So there's two parts to that:\n1) You don't have to convert the weights to probabilities. If you look at the <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\">Pytorch documentation</a> for WeightedRandomSampler:</p>\n\n<blockquote>\n  <p>weights (sequence) – a sequence of weights, not necessary summing up to one</p>\n</blockquote>\n\n<p>If they don't sum to one, they are assumed to be weights and not probabilities, so using the numbers in range 1.0 to 7.74 works fine.</p>\n\n<p>2) How to go from a list of class weights to a list of weights for each image. As you said:</p>\n\n<blockquote>\n  <p>In WeightedRandomSampler we need to provide probabilities for each image.</p>\n</blockquote>\n\n<p>There are different ways you could approach this, but the way I did it (and I think how @Thundo did it as well) was to iterate through all the labels and assign each image the weight of it's least-represented class (or, in other terms, the max weight it could be assigned). The code looks like this:</p>\n\n<p></p><pre><code>def get_sample_weights(labels, weight_per_class):\n    sample_weights = [np.max(np.array(weight_per_class)[np.nonzero(lab)[0]]) for lab in labels]\n    return sample_weights\n</code></pre>\nAnd I call it like this:\n<pre><code>sample_weights = get_sample_weights(labels, np.array([class_weight_log[x] for x in range(28)]))\n</code></pre>\nwhere <code>labels</code> is a numpy array of all the image labels in the training set, and <code>class_weight_log</code> comes from the <code>create_class_weight</code> function here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a><p></p>",
          "rawMarkdown": "So there's two parts to that:\n1) You don't have to convert the weights to probabilities. If you look at the [Pytorch documentation](https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler) for WeightedRandomSampler:\n&gt; weights (sequence) – a sequence of weights, not necessary summing up to one\n\nIf they don't sum to one, they are assumed to be weights and not probabilities, so using the numbers in range 1.0 to 7.74 works fine.\n\n2) How to go from a list of class weights to a list of weights for each image. As you said:\n&gt;  In WeightedRandomSampler we need to provide probabilities for each image.\n\nThere are different ways you could approach this, but the way I did it (and I think how @Thundo did it as well) was to iterate through all the labels and assign each image the weight of it's least-represented class (or, in other terms, the max weight it could be assigned). The code looks like this:\n\n<pre><code>def get_sample_weights(labels, weight_per_class):\n    sample_weights = [np.max(np.array(weight_per_class)[np.nonzero(lab)[0]]) for lab in labels]\n    return sample_weights\n</code></pre>\nAnd I call it like this:\n<pre><code>sample_weights = get_sample_weights(labels, np.array([class_weight_log[x] for x in range(28)]))\n</code></pre>\nwhere `labels` is a numpy array of all the image labels in the training set, and `class_weight_log` comes from the `create_class_weight` function here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065",
          "votes": 2
        }
      ]
    },
    {
      "id": 435750,
      "postDate": "2018-12-08T17:21:44.687Z",
      "content": "<p>Thank you for sharing your insights! I have two questions:\n1) When you say \"multilabel stratification\", do you use the kinds of techniques discussed in this thread? <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819</a> So far I've just been using the \"hacky\" method from one of the notebooks: <code>stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))</code> </p>\n\n<p>2) For weighted sampling, how do calculate the weights for a multilabel problem? I imagine that it might not work well if each combination of labels is given its own weight, like \"0 1 3\" being completely distinct from \"0 1\". When I was thinking about it, I considered assigning the weight of the least represented class. I'm curious if you have a specific approach that you found effective.</p>",
      "rawMarkdown": "Thank you for sharing your insights! I have two questions:\n1) When you say \"multilabel stratification\", do you use the kinds of techniques discussed in this thread? https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819 So far I've just been using the \"hacky\" method from one of the notebooks: `stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))` \n\n2) For weighted sampling, how do calculate the weights for a multilabel problem? I imagine that it might not work well if each combination of labels is given its own weight, like \"0 1 3\" being completely distinct from \"0 1\". When I was thinking about it, I considered assigning the weight of the least represented class. I'm curious if you have a specific approach that you found effective.",
      "votes": 2,
      "replies": [
        {
          "id": 435995,
          "postDate": "2018-12-09T08:42:54.650Z",
          "content": "<p>1) Yes. For reference this is <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">the paper</a> behind the multi-label stratification approach I use.</p>\n\n<pre><code>On the Stratification of Multi-Label Data. Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas, 2011\n</code></pre>\n\n<p>2) I mostly use an heuristic approach. The one you mentioned is my best performing one.</p>",
          "rawMarkdown": "1) Yes. For reference this is [the paper](http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf) behind the multi-label stratification approach I use.\n\n    On the Stratification of Multi-Label Data. Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas, 2011\n\n2) I mostly use an heuristic approach. The one you mentioned is my best performing one.",
          "votes": 2
        },
        {
          "id": 436023,
          "postDate": "2018-12-09T10:41:54.807Z",
          "content": "<p>My view is that work to construct nicely stratified folds in the same proportions of the training set are not worth the effort. I've talked about this in another thread but you will get a more robust estimate of out of sample performance by having 5 folds that are all split roughly in proportion and measuring the variance across folds rather than having 5 validation sets that are all constructed to be similar. I'd be more confident backing models that perform well in this CV framework to generalise better. Of course, stratification is still a reasonable thing to do and it's entirely possible the test set will be in very similar/the same proportions to the training set.</p>\n\n<p>I think weighted sampling or models per class is much more important - I just haven't found anything that's stable at the moment.</p>\n\n<p>Curious to know if anyone if just using BCE as a loss function?</p>",
          "rawMarkdown": "My view is that work to construct nicely stratified folds in the same proportions of the training set are not worth the effort. I've talked about this in another thread but you will get a more robust estimate of out of sample performance by having 5 folds that are all split roughly in proportion and measuring the variance across folds rather than having 5 validation sets that are all constructed to be similar. I'd be more confident backing models that perform well in this CV framework to generalise better. Of course, stratification is still a reasonable thing to do and it's entirely possible the test set will be in very similar/the same proportions to the training set.\n\nI think weighted sampling or models per class is much more important - I just haven't found anything that's stable at the moment.\n\nCurious to know if anyone if just using BCE as a loss function?"
        },
        {
          "id": 436836,
          "postDate": "2018-12-11T02:26:51.287Z",
          "content": "<p>I've used just BCE loss as a loss function, with a weighted sampler. I've also not found anything that stable I feel that my submission was just a matter of luck more-so than anything special that I did. Could you elaborate what you mean by models per class?</p>",
          "rawMarkdown": "I've used just BCE loss as a loss function, with a weighted sampler. I've also not found anything that stable I feel that my submission was just a matter of luck more-so than anything special that I did. Could you elaborate what you mean by models per class?"
        },
        {
          "id": 441853,
          "postDate": "2018-12-19T06:56:59.550Z",
          "content": "<p>recently I'm in trouble，can you tell me how to weighted BCE loss?it just has ''w'' to operate. I also used log bce and operate 'pos_weight'.but the result is bad.i used pytorch</p>",
          "rawMarkdown": "recently I'm in trouble，can you tell me how to weighted BCE loss?it just has ''w'' to operate. I also used log bce and operate 'pos_weight'.but the result is bad.i used pytorch"
        }
      ]
    },
    {
      "id": 446504,
      "postDate": "2018-12-28T07:43:37.737Z",
      "content": "<p>I have tried adding weighted samplers using the methods discussed here but something weird is happening – my training loss decreases much more quickly now but the validation loss is all over the place (high variance), and while I am getting a better local F1 score my LB score is lower. I’m trying to add more regularisation to stop overfitting but it seems like it shouldn’t be happening to being with. Did anyone else have trouble when using it?</p>",
      "rawMarkdown": "I have tried adding weighted samplers using the methods discussed here but something weird is happening – my training loss decreases much more quickly now but the validation loss is all over the place (high variance), and while I am getting a better local F1 score my LB score is lower. I’m trying to add more regularisation to stop overfitting but it seems like it shouldn’t be happening to being with. Did anyone else have trouble when using it?"
    },
    {
      "id": 441069,
      "postDate": "2018-12-18T08:22:54.557Z",
      "content": "<p>do you pretrain models  with external data or you use it all together? \nand do you use all the data since there are labels not present in train dataset and some labels are not certain </p>",
      "rawMarkdown": "do you pretrain models  with external data or you use it all together? \nand do you use all the data since there are labels not present in train dataset and some labels are not certain ",
      "replies": [
        {
          "id": 441997,
          "postDate": "2018-12-19T10:40:57.163Z",
          "content": "<p>I use external+train data for regular training. I did not import samples with uncertain labels.</p>",
          "rawMarkdown": "I use external+train data for regular training. I did not import samples with uncertain labels."
        }
      ]
    },
    {
      "id": 438382,
      "postDate": "2018-12-13T15:11:26.437Z",
      "content": "<p>Thanks  for sharing your experiment.\nwould I ask you a question: do the labels in the your submitted csv have a certain order?</p>",
      "rawMarkdown": "Thanks  for sharing your experiment.\nwould I ask you a question: do the labels in the your submitted csv have a certain order?",
      "replies": [
        {
          "id": 438558,
          "postDate": "2018-12-13T22:03:41.257Z",
          "content": "<p>They're in numerical ascending order in a given row. I don't know if this answers your question...</p>",
          "rawMarkdown": "They're in numerical ascending order in a given row. I don't know if this answers your question..."
        }
      ]
    },
    {
      "id": 438089,
      "postDate": "2018-12-13T04:44:45.997Z",
      "content": "<p>I have used TTA and with average predictions, I was able to boost my public LB score by 0.20. I used single model (No CV ensemble) + BCE + TTA average predictions. No thresholding, No external data.</p>\n\n<p>I tried to use Focal loss and/or thresholding and I couldn't manage to improve my score. </p>\n\n<p>I'm planning to ensemble all 3 individual models that I have trained. </p>",
      "rawMarkdown": "I have used TTA and with average predictions, I was able to boost my public LB score by 0.20. I used single model (No CV ensemble) + BCE + TTA average predictions. No thresholding, No external data.\n\nI tried to use Focal loss and/or thresholding and I couldn't manage to improve my score. \n\nI'm planning to ensemble all 3 individual models that I have trained. ",
      "replies": [
        {
          "id": 438198,
          "postDate": "2018-12-13T08:49:54.320Z",
          "content": "<p>hi, what kind of TTA are you using? I tried flip TTA, but shows no improvement.</p>",
          "rawMarkdown": "hi, what kind of TTA are you using? I tried flip TTA, but shows no improvement."
        },
        {
          "id": 438241,
          "postDate": "2018-12-13T10:40:13.480Z",
          "content": "<p>I'm using flip, rotation, zoom and bightness transformations and then taking mean on my output probabilities.</p>",
          "rawMarkdown": "I'm using flip, rotation, zoom and bightness transformations and then taking mean on my output probabilities.",
          "votes": 1
        },
        {
          "id": 438264,
          "postDate": "2018-12-13T11:34:44.747Z",
          "content": "<p>Hi, could you please provide more facts about your training? Do you use pretrained model? What's size of images? What model do you use?</p>",
          "rawMarkdown": "Hi, could you please provide more facts about your training? Do you use pretrained model? What's size of images? What model do you use?"
        },
        {
          "id": 438299,
          "postDate": "2018-12-13T12:49:27.160Z",
          "content": "<p>thanks for sharing, there might be something wrong with my TTA</p>",
          "rawMarkdown": "thanks for sharing, there might be something wrong with my TTA"
        },
        {
          "id": 438560,
          "postDate": "2018-12-13T22:04:18.337Z",
          "content": "<p>As I said before.. Just added in the last days. It added 0.01 LB.</p>",
          "rawMarkdown": "As I said before.. Just added in the last days. It added 0.01 LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 438071,
      "postDate": "2018-12-13T03:51:24.233Z",
      "content": "<blockquote>\n  <p>\"standard\" augmentation</p>\n</blockquote>\n\n<p>Sorry, I'm a noob. What is considered \"standard\" augmentation for this task?</p>",
      "rawMarkdown": "&gt; \"standard\" augmentation\n\nSorry, I'm a noob. What is considered \"standard\" augmentation for this task?\n",
      "replies": [
        {
          "id": 438077,
          "postDate": "2018-12-13T04:06:30.707Z",
          "content": "<p>I'm using random rotations, zooming, flipping, and brightness. I haven't done rigorous experimenting to see if I need all of those, but it's been working well enough for me so far.</p>",
          "rawMarkdown": "I'm using random rotations, zooming, flipping, and brightness. I haven't done rigorous experimenting to see if I need all of those, but it's been working well enough for me so far."
        },
        {
          "id": 438313,
          "postDate": "2018-12-13T13:15:16.247Z",
          "content": "<p>For me it's flipping, rotation and contrast/brightness</p>",
          "rawMarkdown": "For me it's flipping, rotation and contrast/brightness"
        },
        {
          "id": 438618,
          "postDate": "2018-12-14T00:15:51.667Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!",
          "votes": 1
        },
        {
          "id": 441468,
          "postDate": "2018-12-18T17:26:22.100Z",
          "content": "<p>Are you doing all these augmentations on all the images in a batch? Or is it like sometimes only one of these augmentations is taking place and sometimes the original image(s) is passed?</p>",
          "rawMarkdown": "Are you doing all these augmentations on all the images in a batch? Or is it like sometimes only one of these augmentations is taking place and sometimes the original image(s) is passed?"
        }
      ]
    },
    {
      "id": 437741,
      "postDate": "2018-12-12T12:19:40.730Z",
      "content": "<p>I am bit confused here with the loss function selection. This problem seems to be related to multi class and multi label one and i guess we can't use categorical cross entropy for this.</p>",
      "rawMarkdown": "I am bit confused here with the loss function selection. This problem seems to be related to multi class and multi label one and i guess we can't use categorical cross entropy for this.",
      "replies": [
        {
          "id": 437791,
          "postDate": "2018-12-12T13:52:24.593Z",
          "content": "<p>You should use binary cross entropy instead</p>",
          "rawMarkdown": "You should use binary cross entropy instead",
          "votes": 1
        },
        {
          "id": 437871,
          "postDate": "2018-12-12T17:06:54.807Z",
          "content": "<p>I thought binary cross entropy is for a category class where where one image can have only one class associated. Will this be addressed into binary cross entropy loss function?</p>",
          "rawMarkdown": "I thought binary cross entropy is for a category class where where one image can have only one class associated. Will this be addressed into binary cross entropy loss function?",
          "votes": 1
        },
        {
          "id": 438076,
          "postDate": "2018-12-13T04:02:19.057Z",
          "content": "<p>Binary cross entropy works for data that is multi class and multi label because the problem can be modeled as a binary decision (0 or 1) across <em>n</em> classes (in this case, <em>n</em> is 28). The prediction of one class is independent of another (well maybe not in reality, but we'll ignore the correlations for now), so it makes sense look at each one separately.</p>\n\n<p>For more depth, I found this page helpful: <a href=\"https://gombru.github.io/2018/05/23/cross_entropy_loss/\">https://gombru.github.io/2018/05/23/cross_entropy_loss/</a>.</p>\n\n<p>It says:\n\"Unlike Softmax loss it is independent for each vector component (class), meaning that the loss computed for every vector component is not affected by other component values. That’s why it is used for multi-label classification, were the insight of an element belonging to a certain class should not influence the decision for another class.\"</p>",
          "rawMarkdown": "Binary cross entropy works for data that is multi class and multi label because the problem can be modeled as a binary decision (0 or 1) across _n_ classes (in this case, _n_ is 28). The prediction of one class is independent of another (well maybe not in reality, but we'll ignore the correlations for now), so it makes sense look at each one separately.\n\nFor more depth, I found this page helpful: https://gombru.github.io/2018/05/23/cross_entropy_loss/.\n\nIt says:\n\"Unlike Softmax loss it is independent for each vector component (class), meaning that the loss computed for every vector component is not affected by other component values. That’s why it is used for multi-label classification, were the insight of an element belonging to a certain class should not influence the decision for another class.\""
        }
      ]
    },
    {
      "id": 437195,
      "postDate": "2018-12-11T14:18:59.797Z",
      "content": "<p>Thanks for your share and can you tell me what is the 'flat' threshold ?</p>",
      "rawMarkdown": "Thanks for your share and can you tell me what is the 'flat' threshold ?\n",
      "replies": [
        {
          "id": 437617,
          "postDate": "2018-12-12T07:47:46.483Z",
          "content": "<p>The thresholds are equal for all classes - say, 0.57 or something. Treating all the classes the same.</p>",
          "rawMarkdown": "The thresholds are equal for all classes - say, 0.57 or something. Treating all the classes the same."
        },
        {
          "id": 437657,
          "postDate": "2018-12-12T09:35:10.840Z",
          "content": "<p>Correct... As I said before I'm a bit puzzled by this, but the thresholds scored on the validation set always return a considerably worse LB score.</p>",
          "rawMarkdown": "Correct... As I said before I'm a bit puzzled by this, but the thresholds scored on the validation set always return a considerably worse LB score."
        },
        {
          "id": 437781,
          "postDate": "2018-12-12T13:40:04.330Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 436900,
      "postDate": "2018-12-11T04:56:18.543Z",
      "content": "<p>How many hours are you training for and with which GPU?</p>",
      "rawMarkdown": "How many hours are you training for and with which GPU?",
      "replies": [
        {
          "id": 437661,
          "postDate": "2018-12-12T09:37:37.317Z",
          "content": "<p>With HPA each epoch takes between 40-60 minutes. I train a few epochs for fold.</p>",
          "rawMarkdown": "With HPA each epoch takes between 40-60 minutes. I train a few epochs for fold."
        },
        {
          "id": 440499,
          "postDate": "2018-12-17T16:13:46.933Z",
          "content": "<p>Hi <a href=\"/thundo\">@thundo</a>, <a href=\"/madislemsalu\">@madislemsalu</a>, </p>\n\n<p>Similar question... If I'm running my model with let's say 40k images per epoch, it takes me more than an hour to train an epoch. Lafoss for instance trains his models for 24 epochs so it would take me 24 hours to train a single fold and 8 days for my 8 folds :/</p>\n\n<p>Were you able to converge faster with less epochs ?</p>\n\n<p>Thanks,</p>",
          "rawMarkdown": "Hi @thundo, @madislemsalu, \n\nSimilar question... If I'm running my model with let's say 40k images per epoch, it takes me more than an hour to train an epoch. Lafoss for instance trains his models for 24 epochs so it would take me 24 hours to train a single fold and 8 days for my 8 folds :/\n\nWere you able to converge faster with less epochs ?\n\nThanks,"
        },
        {
          "id": 440662,
          "postDate": "2018-12-17T21:30:55.880Z",
          "content": "<p>Sadly not... </p>\n\n<p>Training is a tradeoff: \n- want to be faster for experiments? use a smaller image size with larger batches, use simpler networks, less epochs, ...\n- want to score big? Do the opposite ;)</p>\n\n<p>There's no free lunch!</p>",
          "rawMarkdown": "Sadly not... \n\nTraining is a tradeoff: \n- want to be faster for experiments? use a smaller image size with larger batches, use simpler networks, less epochs, ...\n- want to score big? Do the opposite ;)\n\nThere's no free lunch!",
          "votes": 2
        }
      ]
    },
    {
      "id": 436771,
      "postDate": "2018-12-10T23:10:22.323Z",
      "content": "<p>Hi Thundo, thanks for sharing your thoughts! I just have couple questions\n1. Can you share how much LB improvements did you get after you train with external data? \n2. Does no tta give you better LB score than using tta? I am using tta right now, and if no tta gives better outcome, i might try dropping this step </p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Hi Thundo, thanks for sharing your thoughts! I just have couple questions\n1. Can you share how much LB improvements did you get after you train with external data? \n2. Does no tta give you better LB score than using tta? I am using tta right now, and if no tta gives better outcome, i might try dropping this step \n\nThank you!",
      "replies": [
        {
          "id": 437118,
          "postDate": "2018-12-11T11:54:20.807Z",
          "content": "<ol>\n<li>No. Using TTA (just flipping) gets a higher score (for me)</li>\n</ol>",
          "rawMarkdown": "2. No. Using TTA (just flipping) gets a higher score (for me)",
          "votes": 2
        },
        {
          "id": 437518,
          "postDate": "2018-12-12T04:34:32.827Z",
          "content": "<p>thank you. I tried TTA with 4 images for each augmentation, and now i want to reduce to just 1 image per each augmentation and see what's the result looks like</p>",
          "rawMarkdown": "thank you. I tried TTA with 4 images for each augmentation, and now i want to reduce to just 1 image per each augmentation and see what's the result looks like"
        },
        {
          "id": 437666,
          "postDate": "2018-12-12T09:40:02.933Z",
          "content": "<ol>\n<li>It's hard to say... My first model with external data was in training during the rescoring, so I don't have an accurate answer, sorry.</li>\n<li>Since I posted this thread I implemented TTA. I got a boost of exactly 0.01.</li>\n</ol>",
          "rawMarkdown": "1. It's hard to say... My first model with external data was in training during the rescoring, so I don't have an accurate answer, sorry.\n2. Since I posted this thread I implemented TTA. I got a boost of exactly 0.01.",
          "votes": 1
        }
      ]
    },
    {
      "id": 436749,
      "postDate": "2018-12-10T21:37:37.087Z",
      "content": "<p>I am currently looking for statistical improvement by simply running models with different random seeds several times (6-10) and combining the results by averaging the predictions before thresholding. 6x might give me 0.01-0.02 LB improvement (big error bar on that statement!). This procedure, compared to single runs, should also provide better judgements of whether proposed improvements are worth keeping (model skill) as the scores are more \"accurate\" .\nI am not doing k-fold CV and I wonder what I am missing. The obvious thing is that I might be more biased in the train/validate split but, with 6-10x random starts, I would think that the bias would not be so bad.</p>\n\n<p>Is there a k-fold advocate who can explain the motivation? I have not found anything on the web that is crystal clear.</p>\n\n<p>Also, for k-fold, I guess that the mean of the individual LB scores is chosen as the \"model skill\" (?) for deciding strategy but that the submission file is constructed with some other procedure that averages the individual predictions somehow. Is this true?</p>",
      "rawMarkdown": "I am currently looking for statistical improvement by simply running models with different random seeds several times (6-10) and combining the results by averaging the predictions before thresholding. 6x might give me 0.01-0.02 LB improvement (big error bar on that statement!). This procedure, compared to single runs, should also provide better judgements of whether proposed improvements are worth keeping (model skill) as the scores are more \"accurate\" .\nI am not doing k-fold CV and I wonder what I am missing. The obvious thing is that I might be more biased in the train/validate split but, with 6-10x random starts, I would think that the bias would not be so bad.\n\nIs there a k-fold advocate who can explain the motivation? I have not found anything on the web that is crystal clear.\n\nAlso, for k-fold, I guess that the mean of the individual LB scores is chosen as the \"model skill\" (?) for deciding strategy but that the submission file is constructed with some other procedure that averages the individual predictions somehow. Is this true?",
      "replies": [
        {
          "id": 436763,
          "postDate": "2018-12-10T22:40:12.890Z",
          "content": "<p>If all you are wanting to do is get an unbiased estimate of out of sample performance then using different seeds for different hold-out sets is fine. However, if you are doing that you may as well then just use k-fold as you can then get the out of fold predictions for ensembling if this is something you plan to do. Of course, you could just average models but there is almost certainly a more optimal way to weight those models - and having the whole of the train data on which to do this (if you train many models over many folds) is more robust. </p>\n\n<p>However, given your LB position you clearly have a strong single model set-up so I wouldn't worry too much. I have briefly tried ensembling model predictions and it wasn't stable though I might be missing something, as hinted at here by current 3rd place, Dmytro: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73934#436430\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73934#436430</a></p>",
          "rawMarkdown": "If all you are wanting to do is get an unbiased estimate of out of sample performance then using different seeds for different hold-out sets is fine. However, if you are doing that you may as well then just use k-fold as you can then get the out of fold predictions for ensembling if this is something you plan to do. Of course, you could just average models but there is almost certainly a more optimal way to weight those models - and having the whole of the train data on which to do this (if you train many models over many folds) is more robust. \n\nHowever, given your LB position you clearly have a strong single model set-up so I wouldn't worry too much. I have briefly tried ensembling model predictions and it wasn't stable though I might be missing something, as hinted at here by current 3rd place, Dmytro: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73934#436430"
        },
        {
          "id": 436766,
          "postDate": "2018-12-10T22:44:42.733Z",
          "content": "<p>lol - I puzzled over that comment too...  Thanks for the info.</p>",
          "rawMarkdown": "lol - I puzzled over that comment too...  Thanks for the info."
        }
      ]
    },
    {
      "id": 436652,
      "postDate": "2018-12-10T17:47:45.990Z",
      "content": "<p>Thanks for your nice sharing!  Just wondering is it possible to train from scratch without pretrained models? \nI tried quite a lot of models like resnet, densenet and inception but none of them get better than 0.4 LB. While my local cv is about 0.48. \nBtw I use BCELoss, all of the above models converge into 0.06</p>",
      "rawMarkdown": "Thanks for your nice sharing!  Just wondering is it possible to train from scratch without pretrained models? \nI tried quite a lot of models like resnet, densenet and inception but none of them get better than 0.4 LB. While my local cv is about 0.48. \nBtw I use BCELoss, all of the above models converge into 0.06\n",
      "replies": [
        {
          "id": 437780,
          "postDate": "2018-12-12T13:38:24.863Z",
          "content": "<p>I did not try many architectures, but I was able to reach more or less the same results with or without pretrained models. Pretrained models usually converge faster...</p>",
          "rawMarkdown": "I did not try many architectures, but I was able to reach more or less the same results with or without pretrained models. Pretrained models usually converge faster..."
        }
      ]
    },
    {
      "id": 435413,
      "postDate": "2018-12-08T02:15:56.863Z",
      "content": "<p>Hill! Thundo,do you used all v1 HPA external data or just a part? </p>",
      "rawMarkdown": "Hill! Thundo,do you used all v1 HPA external data or just a part? ",
      "replies": [
        {
          "id": 435557,
          "postDate": "2018-12-08T08:43:08.533Z",
          "content": "<p>You mean v18? If yes, all excluding the \"uncertain\" samples.</p>",
          "rawMarkdown": "You mean v18? If yes, all excluding the \"uncertain\" samples."
        },
        {
          "id": 435592,
          "postDate": "2018-12-08T10:45:58.547Z",
          "content": "<p>thank you for your answer!</p>",
          "rawMarkdown": "thank you for your answer!",
          "votes": 1
        },
        {
          "id": 436074,
          "postDate": "2018-12-09T14:16:29.117Z",
          "content": "<p>Was there anything that you needed to do for the v18 external data? My model is able to get ~0.67 macro F1 on the original dataset, but with the combined dataset (HPAv18 + competition) I can't get past ~0.45</p>",
          "rawMarkdown": "Was there anything that you needed to do for the v18 external data? My model is able to get ~0.67 macro F1 on the original dataset, but with the combined dataset (HPAv18 + competition) I can't get past ~0.45",
          "votes": 1
        },
        {
          "id": 436335,
          "postDate": "2018-12-10T06:27:16.470Z",
          "content": "<p>get ~0.67 macro F1 on the original dataset? why i can get 0.85 on the valid.is your valid large?<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462</a>  do you kown why?</p>",
          "rawMarkdown": "get ~0.67 macro F1 on the original dataset? why i can get 0.85 on the valid.is your valid large?https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462  do you kown why?"
        }
      ]
    },
    {
      "id": 435019,
      "postDate": "2018-12-07T10:31:19.157Z",
      "content": "<p>Thanks for sharing your ideas.</p>",
      "rawMarkdown": "Thanks for sharing your ideas.",
      "votes": 1
    },
    {
      "id": 434968,
      "postDate": "2018-12-07T08:34:27.840Z",
      "content": "<p>thanks, very helpful!</p>",
      "rawMarkdown": "thanks, very helpful!",
      "votes": 1
    },
    {
      "id": 434947,
      "postDate": "2018-12-07T07:35:32.303Z",
      "content": "<p>Thanks for your sharing! </p>",
      "rawMarkdown": "Thanks for your sharing! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 435859,
      "author_name": "Cozy Doomer",
      "author_url": "",
      "post_date": "2018-12-08T23:11:30.827000",
      "content": "<p>Hey Thundo I found myself in a similar spot! </p>\n\n<ul>\n<li><p>The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.</p></li>\n<li><p>Image size 512 seems to be enough to get a decent score</p></li>\n<li><p>If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)</p></li>\n<li><p>I'm unsure about how the distribution of classes looks right now, as the correction of the leak moved rare cases from private to public (possible overfitting to public leaderboard as there are less rare cases in the private test set)</p></li>\n</ul>\n\n<p>Saying all this I am very much new to deep learning so please correct me if I make wrong assumptions</p>",
      "votes": 5,
      "replies": [
        {
          "id": 435868,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2018-12-08T23:51:25.353000",
          "content": "<p>Hi I was wondering loss functions are you using, and how are you balancing your dataset? I'm using Focal + F1, and balancing by giving each sample a score of 1/proportion_of_class for each positive class in the sample then taking the mean. Though my validation loss is always a lot higher than my training loss, even w/ image augmentation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435884,
          "author_name": "Cozy Doomer",
          "author_url": "",
          "post_date": "2018-12-09T01:31:11.550000",
          "content": "<p>I'm using Focalloss and do balancing mostly with thresholding my output (for each class) </p>\n\n<p>Also using a validation split that has a similar class imbalance helps I think</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435955,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2018-12-09T06:28:07.347000",
          "content": "<p>I see so you don't explicitly do any under/oversampling of each sample?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436000,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-09T08:53:49.907000",
          "content": "<p>&gt; The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.</p>\n\n<p>You're right. Considering download difficulties, increased storage (original and resized versions) and training time, it could be easily an entry barrier for lots of users.</p>\n\n<p>&gt; Image size 512 seems to be enough to get a decent score</p>\n\n<p>I would like to try something bigger but I think I don't have enough resources at my disposal.</p>\n\n<p>&gt; If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)</p>\n\n<p>I tried thresholds based on the predictions of the out-of-fold validation set. The CV score is better but not the LB. Maybe as someone suggested, this is a byproduct of the weighted sampling or maybe just overfitting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436539,
          "author_name": "GhMa",
          "author_url": "",
          "post_date": "2018-12-10T13:52:56.127000",
          "content": "<p>Nice share! \nHmm, Thundo, could you please tell me where I can get those \"external\" data ? \nI haven't read any info about that on the \"Data\" column of the competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436547,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-10T14:06:01.227000",
          "content": "<p>You can find instructions to download the external data here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#432870\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#432870</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434983,
      "author_name": "FlYM",
      "author_url": "",
      "post_date": "2018-12-07T09:05:23.300000",
      "content": "<p>I use TTA (standard), independent thresholds, no-ensemble (at the moment), v1 HPA external data. I think finding the right thresholds is important and for sure would help you. No easy task though.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 435076,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-07T13:10:24.403000",
          "content": "<p>Thanks <a href=\"/arnaurm\">@arnaurm</a>. </p>\n\n<p>I plan to add TTA next. Could you share more about your take on threshold selection?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437696,
          "author_name": "FlYM",
          "author_url": "",
          "post_date": "2018-12-12T10:20:39.800000",
          "content": "<p>I just find the thresholds that fit best with my validation data. Later maybe I will try an ensemble with majority voting. Having said this, don't trust the public leaderboard, I have the feeling, as you said, that some of us are overfitting with the leaked data. I think real score is inbetween 0.59-0.60. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 436979,
      "author_name": "Huang, Shuang",
      "author_url": "",
      "post_date": "2018-12-11T07:11:16.363000",
      "content": "<p>Thanks for your share. I have a question:\nHow many channels do you used, I found I got a higher score with 3 channels(RGB) than with 4 channels(RGBY).</p>",
      "votes": 3,
      "replies": [
        {
          "id": 437658,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T09:35:49.547000",
          "content": "<p>I tried both approaches a while back... In the end I resorted to RGBY.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437690,
          "author_name": "Chris Oosthuizen",
          "author_url": "",
          "post_date": "2018-12-12T10:12:27.290000",
          "content": "<p>The challenge is the the HPA images only has 3 layers. Do you add a yellow layer filled with 0. or null vlaues to make the shape consistent with the competition data images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437707,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T10:45:07.673000",
          "content": "<p>That's what I thought at the beginning... Actually most of them have 4 channels... You can find more explanations and even working code in <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">the external data thread</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 440769,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-12-18T01:04:18.513000",
      "content": "<p>I'm just wondering if anyone has successfully gone over 512 in image size. It looks do-able to me if the batch size is not large. Of course, it'll be very slow...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 440783,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-18T01:17:33.593000",
          "content": "<p>I attempted 1024, but was not able to run enough epochs to make it competitive with my 512 solution. The constraint really is time, and I think most people (who are operating with constrained resources) would be better served continuing to experiment at 512 and building ensembles rather than trying the higher resolution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441641,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-18T22:02:51.457000",
          "content": "<p>I tried something in-between 512 and 1024. Training is painfully slow on my GTX 1070 and it didn't improve score...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446525,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2018-12-28T08:36:24.930000",
          "content": "<p>Batch size is also too small and it actually affects convergence. \nI think that one needs really expensive GPUs to train deep CNNs on 1024x1024 pictures with a decent batch size.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435443,
      "author_name": "Gyeddela",
      "author_url": "",
      "post_date": "2018-12-08T03:35:10.507000",
      "content": "<p>What do you mean by \"cross-validation ensemble\", can you help with some example? Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 435997,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-09T08:43:53.713000",
          "content": "<p>I use a 5-fold CV for training, then I ensemble the predictions of the 5 different models.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436076,
          "author_name": "Gyeddela",
          "author_url": "",
          "post_date": "2018-12-09T14:27:35.033000",
          "content": "<p>Thank You!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436378,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-12-10T07:54:32.830000",
          "content": "<p>hi, how do you ensemble predictions from 5-fold cv models? Do you average all 5 probability on test set? If so, what thresholds do you use? Since 5-fold have differnet best threshold on validation set, it is hard to choose the best threshold on test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436635,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-10T17:05:25.200000",
          "content": "<p>Yes, I use the mean of the predictions. \nThresholds on the out of fold are not so different for me, so I averaged them too.\nBut as I wrote I use a flat threshold for every class.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434834,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-12-07T02:44:43.477000",
      "content": "<p>Thanks for the post and congratulations.\nI use a tilted threshold. Perhaps the flat works for you due to the weighted sampler.</p>\n\n<p>Are you modifying your submission according to the HPA leak? Also, are you training on the images? I am trying that now.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434959,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-07T08:10:21.820000",
          "content": "<p>I tried per-class thresholds using the validation set but got worse results. This could, of course, entirely be related to the public leaderboard distribution.</p>\n\n<p>I use HPA as an extension of my train dataset, so yes, I use it for training. It's hard to say, due the rescoring, how much improved my LB score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435418,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-08T02:28:41.310000",
          "content": "<p>hello! pete,do you mean that you can get so high LB score when you don't use  v1 HPA external data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435876,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-09T00:10:17.107000",
          "content": "<p>I had .504 before the adjustment without HPS usage. After the adjustment, I had .503 (reasonable). Then, I used the HPA data to take advantage of the leak and got 0.55 or so. That was ONLY using the leak csv file posted - none of the images - to fix up the submission file. So far, I have not used the HPA images directly though I plan to do it for training augmentation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436327,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-10T06:07:43.610000",
          "content": "<p>you just used the leak csv file posted to fix up the submission file? is it permitted?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436383,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-10T08:06:47.153000",
          "content": "<p>Yes, just that file. It is permitted. It does not give anyone an advantage because it is public. And it does not appear to detract from the challenge goals. It is worth thinking about what happens on private leaderboard - I have not fully understood the organizer comments on the rare classes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436414,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-10T08:58:00.730000",
          "content": "<p>My understanding is that anyone who can't replicate their public LB score with the leak file by a submission that excludes it is in for a fall on the private LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 437816,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-12T14:53:05.887000",
          "content": "<p>oh,very good , thank you for your useful answer.can you tell me where is the csv file?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434761,
      "author_name": "Renan Rodrigues dos Santos",
      "author_url": "",
      "post_date": "2018-12-06T22:39:26.627000",
      "content": "<p>Calibrating Probability with Undersampling for Unbalanced Classification - <a href=\"https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf\">https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf</a></p>\n\n<p>I'm a noob on this matter but intuitively I think if you oversample (or undersample) your minority (majority) classes, you change their priors. If you choose one sample at random from an unbalanced dataset, the probability of getting a sample from a seldom class is quite low. With weighted sampler, the probability of choosing a sample at random from the same class  increases, thus the need for probability calibration of the network's output when it is trained with an artificially balanced dataset. I might be saying nonsense though. Hopefully, an expert on this subject can comment more on that....\nAnyways, still can't get oversampling to work well in my model</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434966,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-07T08:29:11.453000",
          "content": "<p>Isn't the paper about exclusive classes? We have instead multi-labels. I think this change a bit the context, because the more popular labels don't \"disappear\" when you boost the low-impact ones, they may coexist. </p>\n\n<p>All in all I don't think the distribution changes so dramatically... Bottom line: I don't use any kind of postprocess on the probabilities.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434739,
      "author_name": "Renan Rodrigues dos Santos",
      "author_url": "",
      "post_date": "2018-12-06T21:49:10.043000",
      "content": "<p>@Thundo Thank you for the tips! Do you calibrate your probabilities after using weighted samplers? I tried weighted sampler in my code but got no improvement in public LB (CV score was better though..not sure if I should trust it yet)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 434746,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-06T21:57:53.357000",
          "content": "<p>What do you mean with \"calibrate your probabilities\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440107,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-17T03:40:11.883000",
          "content": "<p>I would suggest trying log dampening on the weights. You can find more info in this thread: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434736,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2018-12-06T21:32:04.793000",
      "content": "<p>Thanks Thundo for sharing those ideas... could you please explain what you call \"weighted samplers\" ? Do you have any kernel that describe this ? Thanks ! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 434738,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-06T21:45:40.993000",
          "content": "<p>Sorry, I don't have a kernel to show... I borrowed the term from PyTorch. I mean that I draw samples for the mini-batches taking class weigths into account...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 434809,
          "author_name": "Mau Dzung",
          "author_url": "",
          "post_date": "2018-12-07T01:08:05.940000",
          "content": "<p>Could you share 'weighted sampler' code in Pytorch for the multi-lables task? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434963,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-07T08:16:46.103000",
          "content": "<p>PyTorch already has a <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\">weighted sampler</a>. What I do is simply set a weight for each training sample. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 436073,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2018-12-09T14:11:44.893000",
          "content": "<p>I was wondering what was the scheme that you used to decide what probability to sample each sample, due to the fact that each sample may have mulitple labels. I took the mean of the sums of the inverse counts for each label in the sample, but it doesn't seem to work very well...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440104,
          "author_name": "Daniel lau",
          "author_url": "",
          "post_date": "2018-12-17T03:36:13.743000",
          "content": "<p>Do you find the best weights sample scheme, I think  per-class weights will not work well, because multi class coexists in one label, if use less class probility will increase large class weights too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443102,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-21T02:45:17.117000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 439577,
      "author_name": "Mouhcine",
      "author_url": "",
      "post_date": "2018-12-15T20:42:57.920000",
      "content": "<p>Nice share !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 439544,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2018-12-15T18:59:51.957000",
      "content": "<p>Using multilabel stratification, external data, and (log-dampened) weighted sampling got me from 0.473 -&gt; 0.494 public LB for my best single model (resnet50), no TTA.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 440111,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2018-12-17T03:52:13.520000",
          "content": "<p>Can you share how to learn this skill - multilabel stratification?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440119,
          "author_name": "Shiv Gowda",
          "author_url": "",
          "post_date": "2018-12-17T04:34:57.797000",
          "content": "<p>You can one-hot encode the labels and call the method in  <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>. The usage is similar to cross validation split in ScikitLearn.  The pip/conda install might not work since the package is dated, just copy over the file (<a href=\"https://github.com/trent-b/iterative-stratification/blob/master/iterstrat/ml_stratifiers.py\">https://github.com/trent-b/iterative-stratification/blob/master/iterstrat/ml_stratifiers.py</a>) to your local workspace and it will work just fine.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 440121,
          "author_name": "Shiv Gowda",
          "author_url": "",
          "post_date": "2018-12-17T04:42:06.457000",
          "content": "<p>The basic premise is that we want to do the train/validation split evenly across all labels. This is important in this competition since we have imbalance in dataset for different labels ( for ex. one end 12000 samples of one label(label=0) and on the other just 11(label=27)). The above package does just that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440132,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2018-12-17T05:35:19.587000",
          "content": "<p>I get it, thank you~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441836,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-19T06:32:29.193000",
          "content": "<p>(log-dampened) weighted sampling？can you share it? i find weighted loss don't work well, the score of middle classes is lowest  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442092,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-19T13:28:33.107000",
          "content": "<p>I used the code from here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442452,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-20T01:26:21.113000",
          "content": "<p>i used the link without work, Maybe I have a problem with my weighting operation. i used pytorch with focal loss.you weight with'w' or 'pos_weight' parameter？ </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442520,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-20T03:57:27.347000",
          "content": "<p>I’ve used weighted sampling, but haven’t tried weighted loss yet. For weighted sampling I found the following thread very useful for Pytorch: <a href=\"https://discuss.pytorch.org/t/balanced-sampling-between-classes-with-torchvision-dataloader/2703/15\">https://discuss.pytorch.org/t/balanced-sampling-between-classes-with-torchvision-dataloader/2703/15</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442621,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-20T08:15:32.957000",
          "content": "<p>thank you .but the link isn't existed.....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442631,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-20T08:40:07.433000",
          "content": "<p><a href=\"/xu666bird\">@xu666bird</a> Remove the full-stop at the end of the link</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 442633,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-20T08:47:31.377000",
          "content": "<p>thank you . it's ok</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442652,
          "author_name": "Criminal Mind",
          "author_url": "",
          "post_date": "2018-12-20T09:25:53.987000",
          "content": "<p><a href=\"/hortonhearsafoo\">@hortonhearsafoo</a> How did you convert your log-dampened weights to probabilities between 0 and 1 for your <code>WeightedRandomSampler</code>  because I see for you are not using log-dampened weights in your loss function and the log-dampened weights are in the range between 1.0 and 7.74 (As per the link you suggested to calculate weights above). My first assumption was you are using these weights in your loss function but I also got the worse results with class weights and Focal Loss (as pointed out by <a href=\"/xu666bird\">@xu666bird</a> ).  In WeightedRanomSampler we need to provide probabilities for each image. Could you please elaborate more on this? Thanks in advance :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442821,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-20T15:15:01.130000",
          "content": "<p>So there's two parts to that:\n1) You don't have to convert the weights to probabilities. If you look at the <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\">Pytorch documentation</a> for WeightedRandomSampler:</p>\n\n<blockquote>\n  <p>weights (sequence) – a sequence of weights, not necessary summing up to one</p>\n</blockquote>\n\n<p>If they don't sum to one, they are assumed to be weights and not probabilities, so using the numbers in range 1.0 to 7.74 works fine.</p>\n\n<p>2) How to go from a list of class weights to a list of weights for each image. As you said:</p>\n\n<blockquote>\n  <p>In WeightedRandomSampler we need to provide probabilities for each image.</p>\n</blockquote>\n\n<p>There are different ways you could approach this, but the way I did it (and I think how @Thundo did it as well) was to iterate through all the labels and assign each image the weight of it's least-represented class (or, in other terms, the max weight it could be assigned). The code looks like this:</p>\n\n<p></p><pre><code>def get_sample_weights(labels, weight_per_class):\n    sample_weights = [np.max(np.array(weight_per_class)[np.nonzero(lab)[0]]) for lab in labels]\n    return sample_weights\n</code></pre>\nAnd I call it like this:\n<pre><code>sample_weights = get_sample_weights(labels, np.array([class_weight_log[x] for x in range(28)]))\n</code></pre>\nwhere <code>labels</code> is a numpy array of all the image labels in the training set, and <code>class_weight_log</code> comes from the <code>create_class_weight</code> function here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a><p></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 435750,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2018-12-08T17:21:44.687000",
      "content": "<p>Thank you for sharing your insights! I have two questions:\n1) When you say \"multilabel stratification\", do you use the kinds of techniques discussed in this thread? <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819</a> So far I've just been using the \"hacky\" method from one of the notebooks: <code>stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))</code> </p>\n\n<p>2) For weighted sampling, how do calculate the weights for a multilabel problem? I imagine that it might not work well if each combination of labels is given its own weight, like \"0 1 3\" being completely distinct from \"0 1\". When I was thinking about it, I considered assigning the weight of the least represented class. I'm curious if you have a specific approach that you found effective.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 435995,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-09T08:42:54.650000",
          "content": "<p>1) Yes. For reference this is <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">the paper</a> behind the multi-label stratification approach I use.</p>\n\n<pre><code>On the Stratification of Multi-Label Data. Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas, 2011\n</code></pre>\n\n<p>2) I mostly use an heuristic approach. The one you mentioned is my best performing one.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 436023,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-09T10:41:54.807000",
          "content": "<p>My view is that work to construct nicely stratified folds in the same proportions of the training set are not worth the effort. I've talked about this in another thread but you will get a more robust estimate of out of sample performance by having 5 folds that are all split roughly in proportion and measuring the variance across folds rather than having 5 validation sets that are all constructed to be similar. I'd be more confident backing models that perform well in this CV framework to generalise better. Of course, stratification is still a reasonable thing to do and it's entirely possible the test set will be in very similar/the same proportions to the training set.</p>\n\n<p>I think weighted sampling or models per class is much more important - I just haven't found anything that's stable at the moment.</p>\n\n<p>Curious to know if anyone if just using BCE as a loss function?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436836,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2018-12-11T02:26:51.287000",
          "content": "<p>I've used just BCE loss as a loss function, with a weighted sampler. I've also not found anything that stable I feel that my submission was just a matter of luck more-so than anything special that I did. Could you elaborate what you mean by models per class?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441853,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-19T06:56:59.550000",
          "content": "<p>recently I'm in trouble，can you tell me how to weighted BCE loss?it just has ''w'' to operate. I also used log bce and operate 'pos_weight'.but the result is bad.i used pytorch</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 446504,
      "author_name": "hedemann",
      "author_url": "",
      "post_date": "2018-12-28T07:43:37.737000",
      "content": "<p>I have tried adding weighted samplers using the methods discussed here but something weird is happening – my training loss decreases much more quickly now but the validation loss is all over the place (high variance), and while I am getting a better local F1 score my LB score is lower. I’m trying to add more regularisation to stop overfitting but it seems like it shouldn’t be happening to being with. Did anyone else have trouble when using it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441069,
      "author_name": "Artyom Lyan",
      "author_url": "",
      "post_date": "2018-12-18T08:22:54.557000",
      "content": "<p>do you pretrain models  with external data or you use it all together? \nand do you use all the data since there are labels not present in train dataset and some labels are not certain </p>",
      "votes": 0,
      "replies": [
        {
          "id": 441997,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-19T10:40:57.163000",
          "content": "<p>I use external+train data for regular training. I did not import samples with uncertain labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 438382,
      "author_name": "Gavroche",
      "author_url": "",
      "post_date": "2018-12-13T15:11:26.437000",
      "content": "<p>Thanks  for sharing your experiment.\nwould I ask you a question: do the labels in the your submitted csv have a certain order?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 438558,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-13T22:03:41.257000",
          "content": "<p>They're in numerical ascending order in a given row. I don't know if this answers your question...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 438089,
      "author_name": "Criminal Mind",
      "author_url": "",
      "post_date": "2018-12-13T04:44:45.997000",
      "content": "<p>I have used TTA and with average predictions, I was able to boost my public LB score by 0.20. I used single model (No CV ensemble) + BCE + TTA average predictions. No thresholding, No external data.</p>\n\n<p>I tried to use Focal loss and/or thresholding and I couldn't manage to improve my score. </p>\n\n<p>I'm planning to ensemble all 3 individual models that I have trained. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 438198,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-12-13T08:49:54.320000",
          "content": "<p>hi, what kind of TTA are you using? I tried flip TTA, but shows no improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438241,
          "author_name": "Criminal Mind",
          "author_url": "",
          "post_date": "2018-12-13T10:40:13.480000",
          "content": "<p>I'm using flip, rotation, zoom and bightness transformations and then taking mean on my output probabilities.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438264,
          "author_name": "Anastasia Smolskaya",
          "author_url": "",
          "post_date": "2018-12-13T11:34:44.747000",
          "content": "<p>Hi, could you please provide more facts about your training? Do you use pretrained model? What's size of images? What model do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438299,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-12-13T12:49:27.160000",
          "content": "<p>thanks for sharing, there might be something wrong with my TTA</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438560,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-13T22:04:18.337000",
          "content": "<p>As I said before.. Just added in the last days. It added 0.01 LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 438071,
      "author_name": "agisga",
      "author_url": "",
      "post_date": "2018-12-13T03:51:24.233000",
      "content": "<blockquote>\n  <p>\"standard\" augmentation</p>\n</blockquote>\n\n<p>Sorry, I'm a noob. What is considered \"standard\" augmentation for this task?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 438077,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-13T04:06:30.707000",
          "content": "<p>I'm using random rotations, zooming, flipping, and brightness. I haven't done rigorous experimenting to see if I need all of those, but it's been working well enough for me so far.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438313,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-13T13:15:16.247000",
          "content": "<p>For me it's flipping, rotation and contrast/brightness</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438618,
          "author_name": "agisga",
          "author_url": "",
          "post_date": "2018-12-14T00:15:51.667000",
          "content": "<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441468,
          "author_name": "Avijit Mitra",
          "author_url": "",
          "post_date": "2018-12-18T17:26:22.100000",
          "content": "<p>Are you doing all these augmentations on all the images in a batch? Or is it like sometimes only one of these augmentations is taking place and sometimes the original image(s) is passed?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437741,
      "author_name": "Amit",
      "author_url": "",
      "post_date": "2018-12-12T12:19:40.730000",
      "content": "<p>I am bit confused here with the loss function selection. This problem seems to be related to multi class and multi label one and i guess we can't use categorical cross entropy for this.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437791,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-12T13:52:24.593000",
          "content": "<p>You should use binary cross entropy instead</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437871,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2018-12-12T17:06:54.807000",
          "content": "<p>I thought binary cross entropy is for a category class where where one image can have only one class associated. Will this be addressed into binary cross entropy loss function?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438076,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2018-12-13T04:02:19.057000",
          "content": "<p>Binary cross entropy works for data that is multi class and multi label because the problem can be modeled as a binary decision (0 or 1) across <em>n</em> classes (in this case, <em>n</em> is 28). The prediction of one class is independent of another (well maybe not in reality, but we'll ignore the correlations for now), so it makes sense look at each one separately.</p>\n\n<p>For more depth, I found this page helpful: <a href=\"https://gombru.github.io/2018/05/23/cross_entropy_loss/\">https://gombru.github.io/2018/05/23/cross_entropy_loss/</a>.</p>\n\n<p>It says:\n\"Unlike Softmax loss it is independent for each vector component (class), meaning that the loss computed for every vector component is not affected by other component values. That’s why it is used for multi-label classification, were the insight of an element belonging to a certain class should not influence the decision for another class.\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437195,
      "author_name": "NullPoint",
      "author_url": "",
      "post_date": "2018-12-11T14:18:59.797000",
      "content": "<p>Thanks for your share and can you tell me what is the 'flat' threshold ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437617,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-12T07:47:46.483000",
          "content": "<p>The thresholds are equal for all classes - say, 0.57 or something. Treating all the classes the same.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437657,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T09:35:10.840000",
          "content": "<p>Correct... As I said before I'm a bit puzzled by this, but the thresholds scored on the validation set always return a considerably worse LB score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437781,
          "author_name": "NullPoint",
          "author_url": "",
          "post_date": "2018-12-12T13:40:04.330000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436900,
      "author_name": "Madis_Lemsalu",
      "author_url": "",
      "post_date": "2018-12-11T04:56:18.543000",
      "content": "<p>How many hours are you training for and with which GPU?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437661,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T09:37:37.317000",
          "content": "<p>With HPA each epoch takes between 40-60 minutes. I train a few epochs for fold.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440499,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-12-17T16:13:46.933000",
          "content": "<p>Hi <a href=\"/thundo\">@thundo</a>, <a href=\"/madislemsalu\">@madislemsalu</a>, </p>\n\n<p>Similar question... If I'm running my model with let's say 40k images per epoch, it takes me more than an hour to train an epoch. Lafoss for instance trains his models for 24 epochs so it would take me 24 hours to train a single fold and 8 days for my 8 folds :/</p>\n\n<p>Were you able to converge faster with less epochs ?</p>\n\n<p>Thanks,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440662,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-17T21:30:55.880000",
          "content": "<p>Sadly not... </p>\n\n<p>Training is a tradeoff: \n- want to be faster for experiments? use a smaller image size with larger batches, use simpler networks, less epochs, ...\n- want to score big? Do the opposite ;)</p>\n\n<p>There's no free lunch!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 436771,
      "author_name": "sx318",
      "author_url": "",
      "post_date": "2018-12-10T23:10:22.323000",
      "content": "<p>Hi Thundo, thanks for sharing your thoughts! I just have couple questions\n1. Can you share how much LB improvements did you get after you train with external data? \n2. Does no tta give you better LB score than using tta? I am using tta right now, and if no tta gives better outcome, i might try dropping this step </p>\n\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437118,
          "author_name": "Mau Dzung",
          "author_url": "",
          "post_date": "2018-12-11T11:54:20.807000",
          "content": "<ol>\n<li>No. Using TTA (just flipping) gets a higher score (for me)</li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 437518,
          "author_name": "sx318",
          "author_url": "",
          "post_date": "2018-12-12T04:34:32.827000",
          "content": "<p>thank you. I tried TTA with 4 images for each augmentation, and now i want to reduce to just 1 image per each augmentation and see what's the result looks like</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437666,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T09:40:02.933000",
          "content": "<ol>\n<li>It's hard to say... My first model with external data was in training during the rescoring, so I don't have an accurate answer, sorry.</li>\n<li>Since I posted this thread I implemented TTA. I got a boost of exactly 0.01.</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 436749,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-12-10T21:37:37.087000",
      "content": "<p>I am currently looking for statistical improvement by simply running models with different random seeds several times (6-10) and combining the results by averaging the predictions before thresholding. 6x might give me 0.01-0.02 LB improvement (big error bar on that statement!). This procedure, compared to single runs, should also provide better judgements of whether proposed improvements are worth keeping (model skill) as the scores are more \"accurate\" .\nI am not doing k-fold CV and I wonder what I am missing. The obvious thing is that I might be more biased in the train/validate split but, with 6-10x random starts, I would think that the bias would not be so bad.</p>\n\n<p>Is there a k-fold advocate who can explain the motivation? I have not found anything on the web that is crystal clear.</p>\n\n<p>Also, for k-fold, I guess that the mean of the individual LB scores is chosen as the \"model skill\" (?) for deciding strategy but that the submission file is constructed with some other procedure that averages the individual predictions somehow. Is this true?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 436763,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-10T22:40:12.890000",
          "content": "<p>If all you are wanting to do is get an unbiased estimate of out of sample performance then using different seeds for different hold-out sets is fine. However, if you are doing that you may as well then just use k-fold as you can then get the out of fold predictions for ensembling if this is something you plan to do. Of course, you could just average models but there is almost certainly a more optimal way to weight those models - and having the whole of the train data on which to do this (if you train many models over many folds) is more robust. </p>\n\n<p>However, given your LB position you clearly have a strong single model set-up so I wouldn't worry too much. I have briefly tried ensembling model predictions and it wasn't stable though I might be missing something, as hinted at here by current 3rd place, Dmytro: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73934#436430\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/73934#436430</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436766,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-10T22:44:42.733000",
          "content": "<p>lol - I puzzled over that comment too...  Thanks for the info.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436652,
      "author_name": "xuan",
      "author_url": "",
      "post_date": "2018-12-10T17:47:45.990000",
      "content": "<p>Thanks for your nice sharing!  Just wondering is it possible to train from scratch without pretrained models? \nI tried quite a lot of models like resnet, densenet and inception but none of them get better than 0.4 LB. While my local cv is about 0.48. \nBtw I use BCELoss, all of the above models converge into 0.06</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437780,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-12T13:38:24.863000",
          "content": "<p>I did not try many architectures, but I was able to reach more or less the same results with or without pretrained models. Pretrained models usually converge faster...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435413,
      "author_name": "ML sutdy ha ha",
      "author_url": "",
      "post_date": "2018-12-08T02:15:56.863000",
      "content": "<p>Hill! Thundo,do you used all v1 HPA external data or just a part? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 435557,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-12-08T08:43:08.533000",
          "content": "<p>You mean v18? If yes, all excluding the \"uncertain\" samples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435592,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-08T10:45:58.547000",
          "content": "<p>thank you for your answer!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436074,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2018-12-09T14:16:29.117000",
          "content": "<p>Was there anything that you needed to do for the v18 external data? My model is able to get ~0.67 macro F1 on the original dataset, but with the combined dataset (HPAv18 + competition) I can't get past ~0.45</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 436335,
          "author_name": "ML sutdy ha ha",
          "author_url": "",
          "post_date": "2018-12-10T06:27:16.470000",
          "content": "<p>get ~0.67 macro F1 on the original dataset? why i can get 0.85 on the valid.is your valid large?<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462</a>  do you kown why?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435019,
      "author_name": "Zhijian Li",
      "author_url": "",
      "post_date": "2018-12-07T10:31:19.157000",
      "content": "<p>Thanks for sharing your ideas.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434968,
      "author_name": "irving",
      "author_url": "",
      "post_date": "2018-12-07T08:34:27.840000",
      "content": "<p>thanks, very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 434947,
      "author_name": "Hogger",
      "author_url": "",
      "post_date": "2018-12-07T07:35:32.303000",
      "content": "<p>Thanks for your sharing! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "434718": "After last rescoring and a couple more submissions I found myself in the top 10. It's not my first competition, but I really wonder how I got so high, even if only temporarily and probably overfitting the public leaderbord. \n\nI want to share what are my best \"selling points\" right now, hoping to get some feedback from others:\n- strong experiment pipeline\n- off-the-shelf pretrained model\n- \"standard\" augmentation\n- HPA external data\n- cross-validation ensemble\n- multilabel stratification\n- weighted samplers\n- lr schedule\n- image size 512\n- \"flat\" threshold\n- no tta\n\nAs I said before, I may be overfitting the public LB badly, so take everything with a pinch of salt....\nThank you!",
    "435859": "Hey Thundo I found myself in a similar spot! \n\n* The key difference of my models with high placement was mostly training with the external data so I think we are kind of being misled as the number of people using it is probably not that high yet.\n\n* Image size 512 seems to be enough to get a decent score\n\n* If you want to improve your score I would say a class dependant threshold is probably a good idea (depends on your model and loss function tough)\n\n* I'm unsure about how the distribution of classes looks right now, as the correction of the leak moved rare cases from private to public (possible overfitting to public leaderboard as there are less rare cases in the private test set)\n\nSaying all this I am very much new to deep learning so please correct me if I make wrong assumptions",
    "434983": "I use TTA (standard), independent thresholds, no-ensemble (at the moment), v1 HPA external data. I think finding the right thresholds is important and for sure would help you. No easy task though.",
    "436979": "Thanks for your share. I have a question:\nHow many channels do you used, I found I got a higher score with 3 channels(RGB) than with 4 channels(RGBY).",
    "440769": "I'm just wondering if anyone has successfully gone over 512 in image size. It looks do-able to me if the batch size is not large. Of course, it'll be very slow...",
    "435443": "What do you mean by \"cross-validation ensemble\", can you help with some example? Thanks!",
    "434834": "Thanks for the post and congratulations.\nI use a tilted threshold. Perhaps the flat works for you due to the weighted sampler.\n\nAre you modifying your submission according to the HPA leak? Also, are you training on the images? I am trying that now.\n",
    "434761": "Calibrating Probability with Undersampling for Unbalanced Classification - https://www3.nd.edu/~dial/publications/dalpozzolo2015calibrating.pdf\n\nI'm a noob on this matter but intuitively I think if you oversample (or undersample) your minority (majority) classes, you change their priors. If you choose one sample at random from an unbalanced dataset, the probability of getting a sample from a seldom class is quite low. With weighted sampler, the probability of choosing a sample at random from the same class  increases, thus the need for probability calibration of the network's output when it is trained with an artificially balanced dataset. I might be saying nonsense though. Hopefully, an expert on this subject can comment more on that....\nAnyways, still can't get oversampling to work well in my model",
    "434739": "@Thundo Thank you for the tips! Do you calibrate your probabilities after using weighted samplers? I tried weighted sampler in my code but got no improvement in public LB (CV score was better though..not sure if I should trust it yet)",
    "434736": "Thanks Thundo for sharing those ideas... could you please explain what you call \"weighted samplers\" ? Do you have any kernel that describe this ? Thanks ! ",
    "439577": "Nice share !",
    "439544": "Using multilabel stratification, external data, and (log-dampened) weighted sampling got me from 0.473 -&gt; 0.494 public LB for my best single model (resnet50), no TTA.",
    "435750": "Thank you for sharing your insights! I have two questions:\n1) When you say \"multilabel stratification\", do you use the kinds of techniques discussed in this thread? https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819 So far I've just been using the \"hacky\" method from one of the notebooks: `stratify = image_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))` \n\n2) For weighted sampling, how do calculate the weights for a multilabel problem? I imagine that it might not work well if each combination of labels is given its own weight, like \"0 1 3\" being completely distinct from \"0 1\". When I was thinking about it, I considered assigning the weight of the least represented class. I'm curious if you have a specific approach that you found effective.",
    "446504": "I have tried adding weighted samplers using the methods discussed here but something weird is happening – my training loss decreases much more quickly now but the validation loss is all over the place (high variance), and while I am getting a better local F1 score my LB score is lower. I’m trying to add more regularisation to stop overfitting but it seems like it shouldn’t be happening to being with. Did anyone else have trouble when using it?",
    "441069": "do you pretrain models  with external data or you use it all together? \nand do you use all the data since there are labels not present in train dataset and some labels are not certain ",
    "438382": "Thanks  for sharing your experiment.\nwould I ask you a question: do the labels in the your submitted csv have a certain order?",
    "438089": "I have used TTA and with average predictions, I was able to boost my public LB score by 0.20. I used single model (No CV ensemble) + BCE + TTA average predictions. No thresholding, No external data.\n\nI tried to use Focal loss and/or thresholding and I couldn't manage to improve my score. \n\nI'm planning to ensemble all 3 individual models that I have trained. ",
    "438071": "&gt; \"standard\" augmentation\n\nSorry, I'm a noob. What is considered \"standard\" augmentation for this task?\n",
    "437741": "I am bit confused here with the loss function selection. This problem seems to be related to multi class and multi label one and i guess we can't use categorical cross entropy for this.",
    "437195": "Thanks for your share and can you tell me what is the 'flat' threshold ?\n",
    "436900": "How many hours are you training for and with which GPU?",
    "436771": "Hi Thundo, thanks for sharing your thoughts! I just have couple questions\n1. Can you share how much LB improvements did you get after you train with external data? \n2. Does no tta give you better LB score than using tta? I am using tta right now, and if no tta gives better outcome, i might try dropping this step \n\nThank you!",
    "436749": "I am currently looking for statistical improvement by simply running models with different random seeds several times (6-10) and combining the results by averaging the predictions before thresholding. 6x might give me 0.01-0.02 LB improvement (big error bar on that statement!). This procedure, compared to single runs, should also provide better judgements of whether proposed improvements are worth keeping (model skill) as the scores are more \"accurate\" .\nI am not doing k-fold CV and I wonder what I am missing. The obvious thing is that I might be more biased in the train/validate split but, with 6-10x random starts, I would think that the bias would not be so bad.\n\nIs there a k-fold advocate who can explain the motivation? I have not found anything on the web that is crystal clear.\n\nAlso, for k-fold, I guess that the mean of the individual LB scores is chosen as the \"model skill\" (?) for deciding strategy but that the submission file is constructed with some other procedure that averages the individual predictions somehow. Is this true?",
    "436652": "Thanks for your nice sharing!  Just wondering is it possible to train from scratch without pretrained models? \nI tried quite a lot of models like resnet, densenet and inception but none of them get better than 0.4 LB. While my local cv is about 0.48. \nBtw I use BCELoss, all of the above models converge into 0.06\n",
    "435413": "Hill! Thundo,do you used all v1 HPA external data or just a part? ",
    "435019": "Thanks for sharing your ideas.",
    "434968": "thanks, very helpful!",
    "434947": "Thanks for your sharing! "
  }
}