{
  "id": 112024,
  "title": "Single model performance",
  "url": "/competitions/understanding_cloud_organization/discussion/112024",
  "author_name": "Psi",
  "post_date": "2019-10-10T12:22:18.768000",
  "votes": 21,
  "comment_count": 61,
  "views": 0,
  "content": "<p>Haven't seen a thread like this yet, but curious about single model CV + LB scores.</p>\n\n<p>I have nothing to report yet (still undecided to join or not).</p>",
  "messages": [
    {
      "id": 645722,
      "postDate": "2019-10-10T12:22:18.770Z",
      "content": "<p>Haven't seen a thread like this yet, but curious about single model CV + LB scores.</p>\n\n<p>I have nothing to report yet (still undecided to join or not).</p>",
      "rawMarkdown": "Haven't seen a thread like this yet, but curious about single model CV + LB scores.\n\nI have nothing to report yet (still undecided to join or not).",
      "votes": 21
    },
    {
      "id": 660328,
      "postDate": "2019-10-29T02:14:43.190Z",
      "content": "<p>segmentation, 1-fold, no tta, 0.664. </p>",
      "rawMarkdown": "segmentation, 1-fold, no tta, 0.664. ",
      "votes": 9,
      "replies": [
        {
          "id": 660332,
          "postDate": "2019-10-29T02:20:45.553Z",
          "content": "<p>Wow, nice job. Have you tried 5 folds yet?</p>",
          "rawMarkdown": "Wow, nice job. Have you tried 5 folds yet?",
          "votes": 1
        },
        {
          "id": 660335,
          "postDate": "2019-10-29T02:27:04.943Z",
          "content": "<p>Based on my experience at the early stage, I would expect ~0.003 boost from 5-fold. </p>",
          "rawMarkdown": "Based on my experience at the early stage, I would expect ~0.003 boost from 5-fold. ",
          "votes": 1
        },
        {
          "id": 660476,
          "postDate": "2019-10-29T07:37:21.297Z",
          "content": "<p>Just segmentation model? Postprocessing?</p>",
          "rawMarkdown": "Just segmentation model? Postprocessing?"
        },
        {
          "id": 660774,
          "postDate": "2019-10-29T15:24:05.377Z",
          "content": "<p>only segmentation, no post-processing.</p>",
          "rawMarkdown": "only segmentation, no post-processing.",
          "votes": 1
        },
        {
          "id": 660788,
          "postDate": "2019-10-29T15:55:24.963Z",
          "content": "<p>That's insane, congratz.</p>",
          "rawMarkdown": "That's insane, congratz.",
          "votes": 1
        },
        {
          "id": 661336,
          "postDate": "2019-10-30T06:46:41.610Z",
          "content": "<p>Thanks, <a href=\"/philippsinger\">@philippsinger</a>. Good luck on the NFL competition! </p>",
          "rawMarkdown": "Thanks, @philippsinger. Good luck on the NFL competition! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 658990,
      "postDate": "2019-10-26T22:06:57.717Z",
      "content": "<p>segmentation, 4fold, no tta, 0.664</p>",
      "rawMarkdown": "segmentation, 4fold, no tta, 0.664",
      "votes": 5,
      "replies": [
        {
          "id": 659061,
          "postDate": "2019-10-27T02:26:05.210Z",
          "content": "<p>hi, phalanx, what about the lb from single fold?</p>",
          "rawMarkdown": "hi, phalanx, what about the lb from single fold?"
        },
        {
          "id": 659064,
          "postDate": "2019-10-27T02:32:47.757Z",
          "content": "<p>How does CV compare with LB. Are training and test similar?</p>",
          "rawMarkdown": "How does CV compare with LB. Are training and test similar?"
        },
        {
          "id": 659076,
          "postDate": "2019-10-27T03:19:25.063Z",
          "content": "<p><a href=\"/garybios\">@garybios</a> \nI didn't confirm  that in LB0.664 model.\nIn my experience, +0.03 by changing single fold to 5 fold average.\n<a href=\"/cdeotte\">@cdeotte</a> \nCV0.653, LB0.664. They don't have a correlation.\nAnnotation method isn' t ideal for me. intersection of all annotator's label is best, not union.</p>",
          "rawMarkdown": "@garybios \nI didn't confirm  that in LB0.664 model.\nIn my experience, +0.03 by changing single fold to 5 fold average.\n@cdeotte \nCV0.653, LB0.664. They don't have a correlation.\nAnnotation method isn' t ideal for me. intersection of all annotator's label is best, not union.",
          "votes": 3
        },
        {
          "id": 659084,
          "postDate": "2019-10-27T03:59:50.073Z",
          "content": "<p><a href=\"/phalanx\">@phalanx</a> oh, are you sure it's +0.03 instead of +0.003?</p>",
          "rawMarkdown": "@phalanx oh, are you sure it's +0.03 instead of +0.003?"
        },
        {
          "id": 659086,
          "postDate": "2019-10-27T04:03:25.387Z",
          "content": "<p>Sorry, 0.003.</p>",
          "rawMarkdown": "Sorry, 0.003."
        },
        {
          "id": 659089,
          "postDate": "2019-10-27T04:08:25.313Z",
          "content": "<p>yep, I got it, thank you.</p>",
          "rawMarkdown": "yep, I got it, thank you."
        },
        {
          "id": 659118,
          "postDate": "2019-10-27T05:52:20.797Z",
          "content": "<p><a href=\"/phalanx\">@phalanx</a>\n are you using your steel competition model?</p>\n\n<p>i haven't download the data yet. is this competition also noisy?</p>",
          "rawMarkdown": "@phalanx\n are you using your steel competition model?\n\ni haven't download the data yet. is this competition also noisy?",
          "votes": 1
        },
        {
          "id": 659119,
          "postDate": "2019-10-27T05:55:26.563Z",
          "content": "<p>Labels are super noisy since they were human annotated without hard proof but just personal preference. </p>",
          "rawMarkdown": "Labels are super noisy since they were human annotated without hard proof but just personal preference. "
        },
        {
          "id": 659173,
          "postDate": "2019-10-27T07:50:03.417Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 659183,
          "postDate": "2019-10-27T08:16:06.737Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 659210,
          "postDate": "2019-10-27T09:19:18.790Z",
          "content": "<p>CV 0.652 LB 0.666</p>\n\n<p>The noisy labels are a big problem, even very similar images can have very different labels. There are some very similar images since we have 2 images per day (Terra and Aqua satellites) and since the train/test split is random some pairs for the same day are split between train and test. I guess organizers should have split the test by time and not randomly. But nevertheless, the labels are so noisy I'm not sure this information leak can lead to any advantage. </p>",
          "rawMarkdown": "CV 0.652 LB 0.666\n\nThe noisy labels are a big problem, even very similar images can have very different labels. There are some very similar images since we have 2 images per day (Terra and Aqua satellites) and since the train/test split is random some pairs for the same day are split between train and test. I guess organizers should have split the test by time and not randomly. But nevertheless, the labels are so noisy I'm not sure this information leak can lead to any advantage. ",
          "votes": 1
        },
        {
          "id": 659336,
          "postDate": "2019-10-27T13:28:33.667Z",
          "content": "<p>I have not looked at data yet, so I'm confused by the comments about different annotators. When I look at the data, will it indicate the name of the annotator? Do images contain annotation information that public EDA notebooks are not showing (like date, location, annotator, etc)?</p>",
          "rawMarkdown": "I have not looked at data yet, so I'm confused by the comments about different annotators. When I look at the data, will it indicate the name of the annotator? Do images contain annotation information that public EDA notebooks are not showing (like date, location, annotator, etc)?"
        },
        {
          "id": 659544,
          "postDate": "2019-10-27T21:46:47.597Z",
          "content": "<p>The labels are crowdsourced by random people through zooniverse. The training they gave this people  was minimal. They just showed them examples of a few of the different cloud types and then showed them some more and asked them to label where they thought the various types were. There was some overlap in that multiple people may have been presented the same image. They took the union of these masks. Interannotator agreement was very low. The masks are basically garbage imo. We dont have any information regarding the annotators we just know there were a few per image. </p>",
          "rawMarkdown": "The labels are crowdsourced by random people through zooniverse. The training they gave this people  was minimal. They just showed them examples of a few of the different cloud types and then showed them some more and asked them to label where they thought the various types were. There was some overlap in that multiple people may have been presented the same image. They took the union of these masks. Interannotator agreement was very low. The masks are basically garbage imo. We dont have any information regarding the annotators we just know there were a few per image. "
        }
      ]
    },
    {
      "id": 646411,
      "postDate": "2019-10-11T08:22:27.607Z",
      "content": "<p>Single model CV .688 / LB .672. No TTA or post processing tried yet. \nEDIT: Bug found. Now CV .65x / LB .67x. Much better.</p>",
      "rawMarkdown": "Single model CV .688 / LB .672. No TTA or post processing tried yet. \nEDIT: Bug found. Now CV .65x / LB .67x. Much better.",
      "votes": 6,
      "replies": [
        {
          "id": 646446,
          "postDate": "2019-10-11T09:00:45.457Z",
          "content": "<p>.688 local CV without any TTA and post processing is quite impressive</p>",
          "rawMarkdown": ".688 local CV without any TTA and post processing is quite impressive"
        },
        {
          "id": 646561,
          "postDate": "2019-10-11T12:30:06.290Z",
          "content": "<p>You have a quite high CV! I don't have much GPU to run 5 folds so, for now, I have LB 0.667 single fold with TTA but results vary quite a lot and are very sensitive to threshold selection. Hopefully with 5 folds they will be more stable.</p>",
          "rawMarkdown": "You have a quite high CV! I don't have much GPU to run 5 folds so, for now, I have LB 0.667 single fold with TTA but results vary quite a lot and are very sensitive to threshold selection. Hopefully with 5 folds they will be more stable."
        },
        {
          "id": 660510,
          "postDate": "2019-10-29T08:34:24.083Z",
          "content": "<p>Segmentation output itself requires postprocessing to produce binary mask, doesn’t it? What does “without postprocessing” mean here?</p>",
          "rawMarkdown": "Segmentation output itself requires postprocessing to produce binary mask, doesn’t it? What does “without postprocessing” mean here?",
          "votes": 1
        }
      ]
    },
    {
      "id": 646173,
      "postDate": "2019-10-10T23:32:27.513Z",
      "content": "<p>single model LB0.656\n5-fold LB0.662</p>\n\n<p>TTA＋post process\nNo classification</p>",
      "rawMarkdown": "single model LB0.656\n5-fold LB0.662\n\nTTA＋post process\nNo classification",
      "votes": 3,
      "replies": [
        {
          "id": 657437,
          "postDate": "2019-10-25T06:26:38.533Z",
          "content": "<p>Are you using qubvel's tta?</p>",
          "rawMarkdown": "Are you using qubvel's tta?",
          "votes": 1
        },
        {
          "id": 659046,
          "postDate": "2019-10-27T01:37:08.210Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes.",
          "votes": 2
        },
        {
          "id": 660267,
          "postDate": "2019-10-29T00:22:27.490Z",
          "content": "<p>when people say k-fold in kaggle, it means run k-fold split data use the same network and hyper-parmas or use the different net / hyper-params?</p>",
          "rawMarkdown": "when people say k-fold in kaggle, it means run k-fold split data use the same network and hyper-parmas or use the different net / hyper-params?"
        },
        {
          "id": 660269,
          "postDate": "2019-10-29T00:24:22.377Z",
          "content": "<p>The same model and hyperparameters, different train and val data.</p>",
          "rawMarkdown": "The same model and hyperparameters, different train and val data.",
          "votes": 1
        },
        {
          "id": 660274,
          "postDate": "2019-10-29T00:29:04.457Z",
          "content": "<p>thanks, so if ensemble n different models ,which means train n*k modes? it's very expensive</p>",
          "rawMarkdown": "thanks, so if ensemble n different models ,which means train n*k modes? it's very expensive"
        },
        {
          "id": 660283,
          "postDate": "2019-10-29T00:39:44.307Z",
          "content": "<p>Yes, it is...</p>",
          "rawMarkdown": "Yes, it is...",
          "votes": 1
        }
      ]
    },
    {
      "id": 645754,
      "postDate": "2019-10-10T13:00:25.150Z",
      "content": "<p>Single model, 75% of training data, no TTA, few basic augmentation, CV 0.642 LB 0.645. The same model, much deeper Unet encoder, much longer training time, 90% training data, valid score not available, LB 0.651. \nCV-LB seems to match, except only a few cases which I am investigating, but in general CV is pretty close to LB, since I have another model with CV 0.626 LB 0.629</p>",
      "rawMarkdown": "Single model, 75% of training data, no TTA, few basic augmentation, CV 0.642 LB 0.645. The same model, much deeper Unet encoder, much longer training time, 90% training data, valid score not available, LB 0.651. \nCV-LB seems to match, except only a few cases which I am investigating, but in general CV is pretty close to LB, since I have another model with CV 0.626 LB 0.629",
      "votes": 3,
      "replies": [
        {
          "id": 645776,
          "postDate": "2019-10-10T13:24:35.760Z",
          "content": "<p>Is there any posprocessing in this model? Btw, it was helpful your advice about dice coef in CV. Thanks</p>",
          "rawMarkdown": "Is there any posprocessing in this model? Btw, it was helpful your advice about dice coef in CV. Thanks"
        },
        {
          "id": 645815,
          "postDate": "2019-10-10T14:05:52.187Z",
          "content": "<p>I made a custom baseline pytorch code which can output validation score after all postprocessing (with all postprocess parameters), then calculating the competition dice metric, and keep track epoch-by-epoch. Since postprocessing is important so I prefer to see its effect while training rather than only doing it after finished training. </p>",
          "rawMarkdown": "I made a custom baseline pytorch code which can output validation score after all postprocessing (with all postprocess parameters), then calculating the competition dice metric, and keep track epoch-by-epoch. Since postprocessing is important so I prefer to see its effect while training rather than only doing it after finished training. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 660841,
      "postDate": "2019-10-29T17:28:26.323Z",
      "content": "<p>Using 3 folds of bounding boxes and no segmentation CV 0.582 and LB 0.611 <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. I'm being careful not to overfit LB this comp :P</p>",
      "rawMarkdown": "Using 3 folds of bounding boxes and no segmentation CV 0.582 and LB 0.611 [here][1]. I'm being careful not to overfit LB this comp :P\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58",
      "votes": 4
    },
    {
      "id": 659757,
      "postDate": "2019-10-28T08:22:23.397Z",
      "content": "<p>3 fold segmentation, threshold search, CV .650~.655, LB .663\nHint: completely same model from 'steel'...</p>",
      "rawMarkdown": "3 fold segmentation, threshold search, CV .650~.655, LB .663\nHint: completely same model from 'steel'...",
      "votes": 1
    },
    {
      "id": 654891,
      "postDate": "2019-10-22T12:52:38.487Z",
      "content": "<p>Unet 0.657 with TTA</p>",
      "rawMarkdown": "Unet 0.657 with TTA",
      "votes": 1,
      "replies": [
        {
          "id": 654906,
          "postDate": "2019-10-22T13:23:03.810Z",
          "content": "<p>What encoder?</p>",
          "rawMarkdown": "What encoder?"
        }
      ]
    },
    {
      "id": 654635,
      "postDate": "2019-10-22T05:53:05.890Z",
      "content": "<p>Basic UNET Resnet 34  20 epochs  CV .58 LB 0.642 , Ño TTA </p>",
      "rawMarkdown": "Basic UNET Resnet 34  20 epochs  CV .58 LB 0.642 , Ño TTA ",
      "votes": 1
    },
    {
      "id": 672749,
      "postDate": "2019-11-14T05:43:50.667Z",
      "content": "<p>I recently got 0.6649 with a single fold + classifier + postprocessing multitask learner network. Am I overfitting or this score is possible without ensembling? </p>",
      "rawMarkdown": "I recently got 0.6649 with a single fold + classifier + postprocessing multitask learner network. Am I overfitting or this score is possible without ensembling? ",
      "votes": 2,
      "replies": [
        {
          "id": 672770,
          "postDate": "2019-11-14T06:23:42.867Z",
          "content": "<p>Congrats, .666 single model is reachable as stated in the forum... I'm still struggling with my .662 segment model.</p>",
          "rawMarkdown": "Congrats, .666 single model is reachable as stated in the forum... I'm still struggling with my .662 segment model."
        },
        {
          "id": 672860,
          "postDate": "2019-11-14T08:30:48.353Z",
          "content": "<p>hi Niu, can I ask what image size you are using?</p>",
          "rawMarkdown": "hi Niu, can I ask what image size you are using?"
        },
        {
          "id": 672995,
          "postDate": "2019-11-14T11:21:10.990Z",
          "content": "<p>I used a lot of sizes, 512x768, 768x1152, 1024x1536, but larger size is not better, not like man</p>",
          "rawMarkdown": "I used a lot of sizes, 512x768, 768x1152, 1024x1536, but larger size is not better, not like man"
        },
        {
          "id": 673044,
          "postDate": "2019-11-14T12:50:58.763Z",
          "content": "<p>I have a single model of  LB 0.6628(CV 0.659), without a classifier and multitask, but with postproc and TTA. I think it just happened to be a good fit for public LB.\nBut if you can do k-fold and each fold will give you around 0.665 LB and similar CV- you might not be overfitting.</p>",
          "rawMarkdown": "I have a single model of  LB 0.6628(CV 0.659), without a classifier and multitask, but with postproc and TTA. I think it just happened to be a good fit for public LB.\nBut if you can do k-fold and each fold will give you around 0.665 LB and similar CV- you might not be overfitting.",
          "votes": 1
        }
      ]
    },
    {
      "id": 646578,
      "postDate": "2019-10-11T12:53:42.520Z",
      "content": "<p>What about image size? Increase size give better results? \n640x960 and 320x480 is giving same score for me.</p>",
      "rawMarkdown": "What about image size? Increase size give better results? \n640x960 and 320x480 is giving same score for me.",
      "votes": 2,
      "replies": [
        {
          "id": 654990,
          "postDate": "2019-10-22T15:03:59.747Z",
          "content": "<p>Have you tried progressive resizing ? i.e train 20 epochs with 320x480 and then load weights , train with 640x960 , while doing so probably change the loss functions , learning rates etc ?</p>\n\n<p>Also why 320x480 ? I saw a few people doing that , I am just getting accustomed to this competition , is there a thread with this rationale ?</p>",
          "rawMarkdown": "Have you tried progressive resizing ? i.e train 20 epochs with 320x480 and then load weights , train with 640x960 , while doing so probably change the loss functions , learning rates etc ?\n\nAlso why 320x480 ? I saw a few people doing that , I am just getting accustomed to this competition , is there a thread with this rationale ?"
        },
        {
          "id": 655381,
          "postDate": "2019-10-23T01:53:03.663Z",
          "content": "<p>I didn't try progressive resizing... 320x480 is the closest to the final size, which serves unet and maintains the ratio 2/3 of the original image</p>",
          "rawMarkdown": "I didn't try progressive resizing... 320x480 is the closest to the final size, which serves unet and maintains the ratio 2/3 of the original image",
          "votes": 1
        }
      ]
    },
    {
      "id": 646237,
      "postDate": "2019-10-11T02:46:44.957Z",
      "content": "<p>Single model 80-20 split : 0.657 CV, 0.653 LB\n5 fold: 0.658 LB</p>",
      "rawMarkdown": "Single model 80-20 split : 0.657 CV, 0.653 LB\n5 fold: 0.658 LB",
      "votes": 2
    },
    {
      "id": 646110,
      "postDate": "2019-10-10T21:49:44.017Z",
      "content": "<p>Single models all in .650-.658 range for both local and LB. Random assortment of models ensembled together, some bad and some good, .663 local and upper .658 on lb. </p>\n\n<p>I'm going to take a pass on this one until the last week or so. I think labels are just too noisy. </p>",
      "rawMarkdown": "Single models all in .650-.658 range for both local and LB. Random assortment of models ensembled together, some bad and some good, .663 local and upper .658 on lb. \n\nI'm going to take a pass on this one until the last week or so. I think labels are just too noisy. ",
      "votes": 2
    },
    {
      "id": 645780,
      "postDate": "2019-10-10T13:27:03.687Z",
      "content": "<p>My best single model:\nkfold 5 CV 0.624 LB 0.626, no TTA, no posprocess\nkfold 5 CV 0.655 LB 0.656 with posprocess\nkfold 5 CV 0.655 LB 0.657 with posprocess + TTA</p>",
      "rawMarkdown": "My best single model:\nkfold 5 CV 0.624 LB 0.626, no TTA, no posprocess\nkfold 5 CV 0.655 LB 0.656 with posprocess\nkfold 5 CV 0.655 LB 0.657 with posprocess + TTA",
      "votes": 2,
      "replies": [
        {
          "id": 654161,
          "postDate": "2019-10-21T14:30:42.137Z",
          "content": "<p>Do you combine TTA、postprocess with the same five models?</p>",
          "rawMarkdown": "Do you combine TTA、postprocess with the same five models?"
        },
        {
          "id": 654258,
          "postDate": "2019-10-21T16:13:02.110Z",
          "content": "<p>I do TTA in each fold, merge and than apply post processing</p>",
          "rawMarkdown": "I do TTA in each fold, merge and than apply post processing"
        },
        {
          "id": 663709,
          "postDate": "2019-11-02T15:05:29.647Z",
          "content": "<p><a href=\"/igormunizims\">@igormunizims</a>  Hi, May I ask a question?  only use posprocess + TTA, your sorce from 0.626 to 0.657, is that right</p>",
          "rawMarkdown": "@igormunizims  Hi, May I ask a question?  only use posprocess + TTA, your sorce from 0.626 to 0.657, is that right"
        },
        {
          "id": 663716,
          "postDate": "2019-11-02T15:20:17.023Z",
          "content": "<p>Hi @hesen. That's right... for my first models post processing increased a lot the score. I have better models now where post processing keeps helping but the gain is lower</p>",
          "rawMarkdown": "Hi @hesen. That's right... for my first models post processing increased a lot the score. I have better models now where post processing keeps helping but the gain is lower"
        },
        {
          "id": 663726,
          "postDate": "2019-11-02T15:41:08.013Z",
          "content": "<p>OK, I see, Thank you for your reply</p>",
          "rawMarkdown": "OK, I see, Thank you for your reply"
        }
      ]
    },
    {
      "id": 660910,
      "postDate": "2019-10-29T19:10:25.940Z",
      "content": "<p>single fold/ out of 10 fold + single model +classification+no TTA   .654</p>",
      "rawMarkdown": "single fold/ out of 10 fold + single model +classification+no TTA   .654",
      "replies": [
        {
          "id": 663420,
          "postDate": "2019-11-02T01:41:59.783Z",
          "content": "<p>what‘ your classifier performance？ I trained a classifier，but avg acc is only 0.77 on my validation data</p>",
          "rawMarkdown": "what‘ your classifier performance？ I trained a classifier，but avg acc is only 0.77 on my validation data"
        },
        {
          "id": 668268,
          "postDate": "2019-11-08T07:48:11.963Z",
          "content": "<p>Me  too.</p>",
          "rawMarkdown": "Me  too."
        }
      ]
    },
    {
      "id": 664040,
      "postDate": "2019-11-03T04:48:24.673Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 646619,
      "postDate": "2019-10-11T13:38:14.510Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 660328,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-10-29T02:14:43.190000",
      "content": "<p>segmentation, 1-fold, no tta, 0.664. </p>",
      "votes": 9,
      "replies": [
        {
          "id": 660332,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-10-29T02:20:45.553000",
          "content": "<p>Wow, nice job. Have you tried 5 folds yet?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660335,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-10-29T02:27:04.943000",
          "content": "<p>Based on my experience at the early stage, I would expect ~0.003 boost from 5-fold. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660476,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-10-29T07:37:21.297000",
          "content": "<p>Just segmentation model? Postprocessing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660774,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-10-29T15:24:05.377000",
          "content": "<p>only segmentation, no post-processing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660788,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-10-29T15:55:24.963000",
          "content": "<p>That's insane, congratz.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 661336,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-10-30T06:46:41.610000",
          "content": "<p>Thanks, <a href=\"/philippsinger\">@philippsinger</a>. Good luck on the NFL competition! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 658990,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2019-10-26T22:06:57.717000",
      "content": "<p>segmentation, 4fold, no tta, 0.664</p>",
      "votes": 5,
      "replies": [
        {
          "id": 659061,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-10-27T02:26:05.210000",
          "content": "<p>hi, phalanx, what about the lb from single fold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659064,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-10-27T02:32:47.757000",
          "content": "<p>How does CV compare with LB. Are training and test similar?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659076,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2019-10-27T03:19:25.063000",
          "content": "<p><a href=\"/garybios\">@garybios</a> \nI didn't confirm  that in LB0.664 model.\nIn my experience, +0.03 by changing single fold to 5 fold average.\n<a href=\"/cdeotte\">@cdeotte</a> \nCV0.653, LB0.664. They don't have a correlation.\nAnnotation method isn' t ideal for me. intersection of all annotator's label is best, not union.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 659084,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-10-27T03:59:50.073000",
          "content": "<p><a href=\"/phalanx\">@phalanx</a> oh, are you sure it's +0.03 instead of +0.003?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659086,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2019-10-27T04:03:25.387000",
          "content": "<p>Sorry, 0.003.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659089,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-10-27T04:08:25.313000",
          "content": "<p>yep, I got it, thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659118,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-10-27T05:52:20.797000",
          "content": "<p><a href=\"/phalanx\">@phalanx</a>\n are you using your steel competition model?</p>\n\n<p>i haven't download the data yet. is this competition also noisy?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 659119,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-10-27T05:55:26.563000",
          "content": "<p>Labels are super noisy since they were human annotated without hard proof but just personal preference. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659173,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-10-27T07:50:03.417000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659183,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-10-27T08:16:06.737000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 659210,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-10-27T09:19:18.790000",
          "content": "<p>CV 0.652 LB 0.666</p>\n\n<p>The noisy labels are a big problem, even very similar images can have very different labels. There are some very similar images since we have 2 images per day (Terra and Aqua satellites) and since the train/test split is random some pairs for the same day are split between train and test. I guess organizers should have split the test by time and not randomly. But nevertheless, the labels are so noisy I'm not sure this information leak can lead to any advantage. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 659336,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-10-27T13:28:33.667000",
          "content": "<p>I have not looked at data yet, so I'm confused by the comments about different annotators. When I look at the data, will it indicate the name of the annotator? Do images contain annotation information that public EDA notebooks are not showing (like date, location, annotator, etc)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659544,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-10-27T21:46:47.597000",
          "content": "<p>The labels are crowdsourced by random people through zooniverse. The training they gave this people  was minimal. They just showed them examples of a few of the different cloud types and then showed them some more and asked them to label where they thought the various types were. There was some overlap in that multiple people may have been presented the same image. They took the union of these masks. Interannotator agreement was very low. The masks are basically garbage imo. We dont have any information regarding the annotators we just know there were a few per image. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 646411,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2019-10-11T08:22:27.607000",
      "content": "<p>Single model CV .688 / LB .672. No TTA or post processing tried yet. \nEDIT: Bug found. Now CV .65x / LB .67x. Much better.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 646446,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-10-11T09:00:45.457000",
          "content": "<p>.688 local CV without any TTA and post processing is quite impressive</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 646561,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-10-11T12:30:06.290000",
          "content": "<p>You have a quite high CV! I don't have much GPU to run 5 folds so, for now, I have LB 0.667 single fold with TTA but results vary quite a lot and are very sensitive to threshold selection. Hopefully with 5 folds they will be more stable.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660510,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-10-29T08:34:24.083000",
          "content": "<p>Segmentation output itself requires postprocessing to produce binary mask, doesn’t it? What does “without postprocessing” mean here?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 646173,
      "author_name": "Yirun Zhang",
      "author_url": "",
      "post_date": "2019-10-10T23:32:27.513000",
      "content": "<p>single model LB0.656\n5-fold LB0.662</p>\n\n<p>TTA＋post process\nNo classification</p>",
      "votes": 3,
      "replies": [
        {
          "id": 657437,
          "author_name": "Cyr1ll",
          "author_url": "",
          "post_date": "2019-10-25T06:26:38.533000",
          "content": "<p>Are you using qubvel's tta?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 659046,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-10-27T01:37:08.210000",
          "content": "<p>Yes.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 660267,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-10-29T00:22:27.490000",
          "content": "<p>when people say k-fold in kaggle, it means run k-fold split data use the same network and hyper-parmas or use the different net / hyper-params?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660269,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-10-29T00:24:22.377000",
          "content": "<p>The same model and hyperparameters, different train and val data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 660274,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-10-29T00:29:04.457000",
          "content": "<p>thanks, so if ensemble n different models ,which means train n*k modes? it's very expensive</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660283,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-10-29T00:39:44.307000",
          "content": "<p>Yes, it is...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 645754,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-10-10T13:00:25.150000",
      "content": "<p>Single model, 75% of training data, no TTA, few basic augmentation, CV 0.642 LB 0.645. The same model, much deeper Unet encoder, much longer training time, 90% training data, valid score not available, LB 0.651. \nCV-LB seems to match, except only a few cases which I am investigating, but in general CV is pretty close to LB, since I have another model with CV 0.626 LB 0.629</p>",
      "votes": 3,
      "replies": [
        {
          "id": 645776,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-10-10T13:24:35.760000",
          "content": "<p>Is there any posprocessing in this model? Btw, it was helpful your advice about dice coef in CV. Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 645815,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-10-10T14:05:52.187000",
          "content": "<p>I made a custom baseline pytorch code which can output validation score after all postprocessing (with all postprocess parameters), then calculating the competition dice metric, and keep track epoch-by-epoch. Since postprocessing is important so I prefer to see its effect while training rather than only doing it after finished training. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 660841,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-10-29T17:28:26.323000",
      "content": "<p>Using 3 folds of bounding boxes and no segmentation CV 0.582 and LB 0.611 <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. I'm being careful not to overfit LB this comp :P</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 659757,
      "author_name": "Endi Niu",
      "author_url": "",
      "post_date": "2019-10-28T08:22:23.397000",
      "content": "<p>3 fold segmentation, threshold search, CV .650~.655, LB .663\nHint: completely same model from 'steel'...</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 654891,
      "author_name": "Shinsei66",
      "author_url": "",
      "post_date": "2019-10-22T12:52:38.487000",
      "content": "<p>Unet 0.657 with TTA</p>",
      "votes": 1,
      "replies": [
        {
          "id": 654906,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-10-22T13:23:03.810000",
          "content": "<p>What encoder?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 654635,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-10-22T05:53:05.890000",
      "content": "<p>Basic UNET Resnet 34  20 epochs  CV .58 LB 0.642 , Ño TTA </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672749,
      "author_name": "0DD1",
      "author_url": "",
      "post_date": "2019-11-14T05:43:50.667000",
      "content": "<p>I recently got 0.6649 with a single fold + classifier + postprocessing multitask learner network. Am I overfitting or this score is possible without ensembling? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 672770,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-14T06:23:42.867000",
          "content": "<p>Congrats, .666 single model is reachable as stated in the forum... I'm still struggling with my .662 segment model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672860,
          "author_name": "YoungLamb",
          "author_url": "",
          "post_date": "2019-11-14T08:30:48.353000",
          "content": "<p>hi Niu, can I ask what image size you are using?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672995,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-14T11:21:10.990000",
          "content": "<p>I used a lot of sizes, 512x768, 768x1152, 1024x1536, but larger size is not better, not like man</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673044,
          "author_name": "Victor Zaguskin",
          "author_url": "",
          "post_date": "2019-11-14T12:50:58.763000",
          "content": "<p>I have a single model of  LB 0.6628(CV 0.659), without a classifier and multitask, but with postproc and TTA. I think it just happened to be a good fit for public LB.\nBut if you can do k-fold and each fold will give you around 0.665 LB and similar CV- you might not be overfitting.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 646578,
      "author_name": "IgorMuniz",
      "author_url": "",
      "post_date": "2019-10-11T12:53:42.520000",
      "content": "<p>What about image size? Increase size give better results? \n640x960 and 320x480 is giving same score for me.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 654990,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2019-10-22T15:03:59.747000",
          "content": "<p>Have you tried progressive resizing ? i.e train 20 epochs with 320x480 and then load weights , train with 640x960 , while doing so probably change the loss functions , learning rates etc ?</p>\n\n<p>Also why 320x480 ? I saw a few people doing that , I am just getting accustomed to this competition , is there a thread with this rationale ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655381,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-10-23T01:53:03.663000",
          "content": "<p>I didn't try progressive resizing... 320x480 is the closest to the final size, which serves unet and maintains the ratio 2/3 of the original image</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 646237,
      "author_name": "0DD1",
      "author_url": "",
      "post_date": "2019-10-11T02:46:44.957000",
      "content": "<p>Single model 80-20 split : 0.657 CV, 0.653 LB\n5 fold: 0.658 LB</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 646110,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-10-10T21:49:44.017000",
      "content": "<p>Single models all in .650-.658 range for both local and LB. Random assortment of models ensembled together, some bad and some good, .663 local and upper .658 on lb. </p>\n\n<p>I'm going to take a pass on this one until the last week or so. I think labels are just too noisy. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 645780,
      "author_name": "IgorMuniz",
      "author_url": "",
      "post_date": "2019-10-10T13:27:03.687000",
      "content": "<p>My best single model:\nkfold 5 CV 0.624 LB 0.626, no TTA, no posprocess\nkfold 5 CV 0.655 LB 0.656 with posprocess\nkfold 5 CV 0.655 LB 0.657 with posprocess + TTA</p>",
      "votes": 2,
      "replies": [
        {
          "id": 654161,
          "author_name": "cowarder",
          "author_url": "",
          "post_date": "2019-10-21T14:30:42.137000",
          "content": "<p>Do you combine TTA、postprocess with the same five models?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 654258,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-10-21T16:13:02.110000",
          "content": "<p>I do TTA in each fold, merge and than apply post processing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 663709,
          "author_name": "He",
          "author_url": "",
          "post_date": "2019-11-02T15:05:29.647000",
          "content": "<p><a href=\"/igormunizims\">@igormunizims</a>  Hi, May I ask a question?  only use posprocess + TTA, your sorce from 0.626 to 0.657, is that right</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 663716,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-11-02T15:20:17.023000",
          "content": "<p>Hi @hesen. That's right... for my first models post processing increased a lot the score. I have better models now where post processing keeps helping but the gain is lower</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 663726,
          "author_name": "He",
          "author_url": "",
          "post_date": "2019-11-02T15:41:08.013000",
          "content": "<p>OK, I see, Thank you for your reply</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 660910,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-10-29T19:10:25.940000",
      "content": "<p>single fold/ out of 10 fold + single model +classification+no TTA   .654</p>",
      "votes": 0,
      "replies": [
        {
          "id": 663420,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-11-02T01:41:59.783000",
          "content": "<p>what‘ your classifier performance？ I trained a classifier，but avg acc is only 0.77 on my validation data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668268,
          "author_name": "哈尔的移动城堡",
          "author_url": "",
          "post_date": "2019-11-08T07:48:11.963000",
          "content": "<p>Me  too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 664040,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-03T04:48:24.673000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 646619,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-11T13:38:14.510000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "645722": "Haven't seen a thread like this yet, but curious about single model CV + LB scores.\n\nI have nothing to report yet (still undecided to join or not).",
    "660328": "segmentation, 1-fold, no tta, 0.664. ",
    "658990": "segmentation, 4fold, no tta, 0.664",
    "646411": "Single model CV .688 / LB .672. No TTA or post processing tried yet. \nEDIT: Bug found. Now CV .65x / LB .67x. Much better.",
    "646173": "single model LB0.656\n5-fold LB0.662\n\nTTA＋post process\nNo classification",
    "645754": "Single model, 75% of training data, no TTA, few basic augmentation, CV 0.642 LB 0.645. The same model, much deeper Unet encoder, much longer training time, 90% training data, valid score not available, LB 0.651. \nCV-LB seems to match, except only a few cases which I am investigating, but in general CV is pretty close to LB, since I have another model with CV 0.626 LB 0.629",
    "660841": "Using 3 folds of bounding boxes and no segmentation CV 0.582 and LB 0.611 [here][1]. I'm being careful not to overfit LB this comp :P\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58",
    "659757": "3 fold segmentation, threshold search, CV .650~.655, LB .663\nHint: completely same model from 'steel'...",
    "654891": "Unet 0.657 with TTA",
    "654635": "Basic UNET Resnet 34  20 epochs  CV .58 LB 0.642 , Ño TTA ",
    "672749": "I recently got 0.6649 with a single fold + classifier + postprocessing multitask learner network. Am I overfitting or this score is possible without ensembling? ",
    "646578": "What about image size? Increase size give better results? \n640x960 and 320x480 is giving same score for me.",
    "646237": "Single model 80-20 split : 0.657 CV, 0.653 LB\n5 fold: 0.658 LB",
    "646110": "Single models all in .650-.658 range for both local and LB. Random assortment of models ensembled together, some bad and some good, .663 local and upper .658 on lb. \n\nI'm going to take a pass on this one until the last week or so. I think labels are just too noisy. ",
    "645780": "My best single model:\nkfold 5 CV 0.624 LB 0.626, no TTA, no posprocess\nkfold 5 CV 0.655 LB 0.656 with posprocess\nkfold 5 CV 0.655 LB 0.657 with posprocess + TTA",
    "660910": "single fold/ out of 10 fold + single model +classification+no TTA   .654",
    "664040": "",
    "646619": ""
  }
}