{
  "id": 97860,
  "title": "CV vs LB Scores",
  "url": "/competitions/aptos2019-blindness-detection/discussion/97860",
  "author_name": "Abhishek Thakur",
  "post_date": "2019-06-29T09:27:16.181000",
  "votes": 63,
  "comment_count": 139,
  "views": 0,
  "content": "<p>Let’s discuss CV and LB Scores here.</p>\n\n<p>To start with, My CV: 0.90, LB: 0.745</p>\n\n<p>So, it seems like the distribution of test data is not same as training data. What are your CV and LB scores ? ;)</p>",
  "messages": [
    {
      "id": 564326,
      "postDate": "2019-06-29T09:27:16.180Z",
      "content": "<p>Let’s discuss CV and LB Scores here.</p>\n\n<p>To start with, My CV: 0.90, LB: 0.745</p>\n\n<p>So, it seems like the distribution of test data is not same as training data. What are your CV and LB scores ? ;)</p>",
      "rawMarkdown": "Let’s discuss CV and LB Scores here.\n\nTo start with, My CV: 0.90, LB: 0.745\n\nSo, it seems like the distribution of test data is not same as training data. What are your CV and LB scores ? ;)",
      "votes": 61
    },
    {
      "id": 574045,
      "postDate": "2019-07-13T06:42:09.147Z",
      "content": "<p>5-fold CV: 0.8954\nLB: 80.0</p>\n\n<p>resnext101_32x16d /w pre-training on previous competition data and no TTA\ntreating as regression problem</p>",
      "rawMarkdown": "5-fold CV: 0.8954\nLB: 80.0\n\nresnext101_32x16d /w pre-training on previous competition data and no TTA\ntreating as regression problem",
      "votes": 22,
      "replies": [
        {
          "id": 575339,
          "postDate": "2019-07-15T09:56:57.833Z",
          "content": "<p>Hi Tom,just wanted to check how much time it took to train the model?</p>",
          "rawMarkdown": "Hi Tom,just wanted to check how much time it took to train the model?"
        },
        {
          "id": 575567,
          "postDate": "2019-07-15T16:46:30.567Z",
          "content": "<p>Hi GSD,</p>\n\n<p>img size = 256\nbatch size = 32\ngradient accumulation steps = 2\nmixed precision = True</p>\n\n<p>takes about 1.2 minute per epoch on RTX 2080ti (approx. 3000 images).</p>",
          "rawMarkdown": "Hi GSD,\n\nimg size = 256\nbatch size = 32\ngradient accumulation steps = 2\nmixed precision = True\n\ntakes about 1.2 minute per epoch on RTX 2080ti (approx. 3000 images).",
          "votes": 4
        },
        {
          "id": 575599,
          "postDate": "2019-07-15T17:54:26.310Z",
          "content": "<p>Did you finetune with imagenet weights or from scratch?</p>",
          "rawMarkdown": "Did you finetune with imagenet weights or from scratch?"
        },
        {
          "id": 575633,
          "postDate": "2019-07-15T18:41:09.090Z",
          "content": "<p>I'm using the pre-trained instagram weights which can be found here: <a href=\"https://github.com/facebookresearch/WSL-Images\">https://github.com/facebookresearch/WSL-Images</a>.</p>\n\n<p>Very simple training procedure for the 0.80:</p>\n\n<p>1) Load pre-trained instagram weights for 32x16d model\n2) Fine tune on old competition data with early stopping based on current competition data\n3) Fine tune for small number of epoch on new competition data</p>\n\n<p>BTW congrats on your triple GM, I enjoyed the youtube video :)</p>",
          "rawMarkdown": "I'm using the pre-trained instagram weights which can be found here: https://github.com/facebookresearch/WSL-Images.\n\nVery simple training procedure for the 0.80:\n\n1) Load pre-trained instagram weights for 32x16d model\n2) Fine tune on old competition data with early stopping based on current competition data\n3) Fine tune for small number of epoch on new competition data\n\nBTW congrats on your triple GM, I enjoyed the youtube video :)",
          "votes": 23
        },
        {
          "id": 576726,
          "postDate": "2019-07-15T23:55:13.517Z",
          "content": "<p>Hi Tom, can you tell my about what \"early stopping based on current competition data\" means? I'm trying to pretrain with previous competition data. thanks for your sharing!</p>",
          "rawMarkdown": "Hi Tom, can you tell my about what \"early stopping based on current competition data\" means? I'm trying to pretrain with previous competition data. thanks for your sharing!"
        },
        {
          "id": 576772,
          "postDate": "2019-07-16T01:57:00.767Z",
          "content": "<p>Maybe while pretraining, use this new train set as validation set.</p>",
          "rawMarkdown": "Maybe while pretraining, use this new train set as validation set."
        },
        {
          "id": 576800,
          "postDate": "2019-07-16T03:13:52.520Z",
          "content": "<p>Thanks Tom..Much helpful</p>",
          "rawMarkdown": "Thanks Tom..Much helpful"
        },
        {
          "id": 576820,
          "postDate": "2019-07-16T04:05:15.373Z",
          "content": "<p><a href=\"/jionie\">@jionie</a> oh I get it. thanks for the reply!</p>",
          "rawMarkdown": "@jionie oh I get it. thanks for the reply!"
        },
        {
          "id": 576898,
          "postDate": "2019-07-16T06:42:21.027Z",
          "content": "<p>Thanks Tom for ur insight!! U mentioned that u r treating it as regression problem...Is it normal(simple) Regression or Ordinal Regression?</p>",
          "rawMarkdown": "Thanks Tom for ur insight!! U mentioned that u r treating it as regression problem...Is it normal(simple) Regression or Ordinal Regression?"
        },
        {
          "id": 576913,
          "postDate": "2019-07-16T06:59:12.610Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  I just used simple regression with mse loss (same as Abhishek's kernel).\n<a href=\"/yangsaewon\">@yangsaewon</a> jionie is correct, i used the new data as a validation set for the old for the initial training.</p>\n\n<p>I've added a kernel using the pre-trained instagram weights if anyone is interested:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/instagram-to-aptos-resnext-101-32x8d\">https://www.kaggle.com/taindow/instagram-to-aptos-resnext-101-32x8d</a></p>",
          "rawMarkdown": "@bibek777  I just used simple regression with mse loss (same as Abhishek's kernel).\n@yangsaewon jionie is correct, i used the new data as a validation set for the old for the initial training.\n\nI've added a kernel using the pre-trained instagram weights if anyone is interested:\n\nhttps://www.kaggle.com/taindow/instagram-to-aptos-resnext-101-32x8d",
          "votes": 3
        },
        {
          "id": 576932,
          "postDate": "2019-07-16T07:16:16.470Z",
          "content": "<p>Are you guys doing early stopping with loss or kappa btw?</p>",
          "rawMarkdown": "Are you guys doing early stopping with loss or kappa btw?"
        },
        {
          "id": 577868,
          "postDate": "2019-07-17T06:35:28.627Z",
          "content": "<p>im early stopping by kappa but its not of any use it seems.</p>",
          "rawMarkdown": "im early stopping by kappa but its not of any use it seems."
        },
        {
          "id": 577975,
          "postDate": "2019-07-17T08:44:07.033Z",
          "content": "<p>From my experience, early stopping with Kappa is quite unstable. Seems to me that using MSE works better.</p>",
          "rawMarkdown": "From my experience, early stopping with Kappa is quite unstable. Seems to me that using MSE works better."
        },
        {
          "id": 578148,
          "postDate": "2019-07-17T11:47:28.863Z",
          "content": "<p>Yeah I've been using kappa for early stopping, but agree with others it seems relatively unstable.</p>\n\n<p>As a general update, increasingly local cv scores of &gt;0.90 have resulted in lower LB scores between 0.77 and 0.80. </p>",
          "rawMarkdown": "Yeah I've been using kappa for early stopping, but agree with others it seems relatively unstable.\n\nAs a general update, increasingly local cv scores of &gt;0.90 have resulted in lower LB scores between 0.77 and 0.80. ",
          "votes": 1
        },
        {
          "id": 607459,
          "postDate": "2019-08-25T09:41:02.633Z",
          "content": "<p>@ tom while doing pre train using old data\n1) did you take all data or subset of it  so that ratio of all severity levels more or less similar\n2) how much kp score u got using old data on  Validation set</p>",
          "rawMarkdown": "@ tom while doing pre train using old data\n1) did you take all data or subset of it  so that ratio of all severity levels more or less similar\n2) how much kp score u got using old data on  Validation set"
        }
      ]
    },
    {
      "id": 565061,
      "postDate": "2019-06-30T11:26:22.727Z",
      "content": "<p>I think the obvious difference of train and test data is the distribution of image size. <br>\nAnd more to say, the image size of the less risk retinopathy might be smaller.  </p>\n\n<p><strong>train image size</strong> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F6beafd10871da3a7b80595b874203218%2Ftrain_img_size.png?generation=1561893474326004&amp;alt=media\" alt=\"train image size\"></p>\n\n<p><strong>test image size</strong> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F64ae4c8e765fdba37d1dd25b9a3dff96%2Ftest_img_size.png?generation=1561893554727501&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I think the obvious difference of train and test data is the distribution of image size.  \nAnd more to say, the image size of the less risk retinopathy might be smaller.  \n\n**train image size**  \n![train image size](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F6beafd10871da3a7b80595b874203218%2Ftrain_img_size.png?generation=1561893474326004&amp;alt=media)\n\n**test image size**  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F64ae4c8e765fdba37d1dd25b9a3dff96%2Ftest_img_size.png?generation=1561893554727501&amp;alt=media)\n\n",
      "votes": 21,
      "replies": [
        {
          "id": 565401,
          "postDate": "2019-06-30T22:43:00.533Z",
          "content": "<p>Agreed. I've looked at (actually looked at) a sample of the test set. And from that sample, it looks true. I plan on looking at some simple statistics and checking PSI later to get a better idea about what has changed exactly.</p>",
          "rawMarkdown": "Agreed. I've looked at (actually looked at) a sample of the test set. And from that sample, it looks true. I plan on looking at some simple statistics and checking PSI later to get a better idea about what has changed exactly."
        },
        {
          "id": 565422,
          "postDate": "2019-06-30T23:37:16.793Z",
          "content": "<p>Yes.  </p>\n\n<p>&gt; And more to say, the image size of the less risk retinopathy might be smaller.  </p>\n\n<p>My above comment is from this pictures.  </p>\n\n<p><strong>Diagnosis vs Image Row Size</strong> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F0578f9965b6f809b6d5ae27da0f7f0f2%2Frow.png?generation=1561937448006275&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Diagnosis vs Image Column Size</strong> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5be6f48f77a2d77ea2ba1b0f4c4553d7%2Fcol.png?generation=1561937472495149&amp;alt=media\" alt=\"\"></p>\n\n<p>I think the way to resize and crop images would be the key to tackle this kind of problems ( TTA is also important ). <br>\nAnd more, 5th position solution in the previous competition had used the information of image size. But I'm not sure that solution was safe or not in the view of overfitting.</p>",
          "rawMarkdown": "Yes.  \n\n&gt; And more to say, the image size of the less risk retinopathy might be smaller.  \n\nMy above comment is from this pictures.  \n\n**Diagnosis vs Image Row Size** \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F0578f9965b6f809b6d5ae27da0f7f0f2%2Frow.png?generation=1561937448006275&amp;alt=media)\n\n**Diagnosis vs Image Column Size** \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5be6f48f77a2d77ea2ba1b0f4c4553d7%2Fcol.png?generation=1561937472495149&amp;alt=media)\n\n\nI think the way to resize and crop images would be the key to tackle this kind of problems ( TTA is also important ).  \nAnd more, 5th position solution in the previous competition had used the information of image size. But I'm not sure that solution was safe or not in the view of overfitting.",
          "votes": 2
        },
        {
          "id": 565519,
          "postDate": "2019-07-01T03:51:49.867Z",
          "content": "<p>This is what the predicted distribution on the test set looks like for me:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media\" alt=\"\"></p>\n\n<p>It looks very different than the predicted values of the training set. So either I have lots of overfitting --which I probably do-- or the distributions are quite different.</p>",
          "rawMarkdown": "This is what the predicted distribution on the test set looks like for me:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media)\n\nIt looks very different than the predicted values of the training set. So either I have lots of overfitting --which I probably do-- or the distributions are quite different.",
          "votes": 5
        },
        {
          "id": 565524,
          "postDate": "2019-07-01T04:27:17.953Z",
          "content": "<p><a href=\"/puremath86\">@puremath86</a> \nThanks for sharing your inference. <br>\nIt looks you did not optimize your predicted values  with QWK. <br>\nIt is interesting for me to see the histogram after you applied optimization.</p>",
          "rawMarkdown": "@puremath86 \nThanks for sharing your inference.  \nIt looks you did not optimize your predicted values  with QWK.  \nIt is interesting for me to see the histogram after you applied optimization."
        },
        {
          "id": 565526,
          "postDate": "2019-07-01T04:43:54.423Z",
          "content": "<p>I plan on treating the problem as a classification problem in the future --but for now I am treating it as a regression. Hopefully, the residual analysis will be enlightening. </p>",
          "rawMarkdown": "I plan on treating the problem as a classification problem in the future --but for now I am treating it as a regression. Hopefully, the residual analysis will be enlightening. "
        },
        {
          "id": 565562,
          "postDate": "2019-07-01T05:44:24.183Z",
          "content": "<p>In <a href=\"https://www.kaggle.com/miklgr500/auto-encoder\">my research</a> with autoencoder I obtain the similar results, but not so clear for interpretation.</p>",
          "rawMarkdown": "In [my research](https://www.kaggle.com/miklgr500/auto-encoder) with autoencoder I obtain the similar results, but not so clear for interpretation.",
          "votes": 1
        },
        {
          "id": 565564,
          "postDate": "2019-07-01T05:50:10.513Z",
          "content": "<p>Those are some interesting embeddings it learned. Two distinct groups of NR. I wonder why?</p>",
          "rawMarkdown": "Those are some interesting embeddings it learned. Two distinct groups of NR. I wonder why?"
        },
        {
          "id": 565568,
          "postDate": "2019-07-01T06:00:23.100Z",
          "content": "<p>All people have two eyes:)</p>",
          "rawMarkdown": "All people have two eyes:)",
          "votes": 2
        },
        {
          "id": 565574,
          "postDate": "2019-07-01T06:11:32.537Z",
          "content": "<p>Is that really the reason? That make's sense. And not all people have two eyes...<img src=\"https://i.pinimg.com/236x/57/d3/e7/57d3e7f7603d3f5adf2d64aca06df5ed.jpg\" alt=\"\"></p>",
          "rawMarkdown": "Is that really the reason? That make's sense. And not all people have two eyes...![](https://i.pinimg.com/236x/57/d3/e7/57d3e7f7603d3f5adf2d64aca06df5ed.jpg)",
          "votes": 6
        },
        {
          "id": 565603,
          "postDate": "2019-07-01T06:50:45.463Z",
          "content": "<p>Lol :)\nSorry, I'm not formulate my idea very well.\nYes, the right and left eyes are different; they have different positions of the nerve and blood vessels.\nBut in the present state of my research, this is just a hypothesis that an autoencoder can separate two types of eyes. Maybe if I used the introduction of a loss of triplet, they would be the same, but then I influence the model with existing labels and can not make an independent conclusion with the distribution of labels about the data set.</p>",
          "rawMarkdown": "Lol :)\nSorry, I'm not formulate my idea very well.\nYes, the right and left eyes are different; they have different positions of the nerve and blood vessels.\nBut in the present state of my research, this is just a hypothesis that an autoencoder can separate two types of eyes. Maybe if I used the introduction of a loss of triplet, they would be the same, but then I influence the model with existing labels and can not make an independent conclusion with the distribution of labels about the data set.",
          "votes": 1
        },
        {
          "id": 566319,
          "postDate": "2019-07-02T03:50:58.040Z",
          "content": "<p>I don't know that it is helpful to be able to distinguish a left eye from a right eye --especially with all of the data augmentation you'll end up using. </p>",
          "rawMarkdown": "I don't know that it is helpful to be able to distinguish a left eye from a right eye --especially with all of the data augmentation you'll end up using. "
        }
      ]
    },
    {
      "id": 579277,
      "postDate": "2019-07-18T17:09:09.960Z",
      "content": "<p>This is very strange. Using exactly the same cv fold splits each time, I've had the following results:</p>\n\n<p>|Model   |fold0|    fold1|  fold2|  fold3|  fold4|  CV| LB|\n| --- | --- | --- | --- | --- | --- |\nModel 0 |0.8970 |0.8983|    0.9024  |0.8920 |0.8874 |0.8954 |0.80|\nModel 1 |0.9015|    0.9052| 0.9085| 0.9049  |0.8869|    0.9014| 0.77|\nModel 2 |0.9251|    0.9158| 0.9288| 0.9180| 0.9168| 0.9209  |0.79|</p>\n\n<p>The  training set has 3662 pictures, so each validation fold has approx. 730 images. The test data has 1928 images. Yet we're seeing way more unstable results on the test data than we are across folds. Something is very different about the test data (which people have already noted, for sure).</p>\n\n<p>Even when you eyeball it (😒 ) you can see a pronounced difference between train and test:</p>\n\n<p><strong>Train</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2F1d6b371c0112d928a782b52b2f265287%2Frsz_screenshot_from_2019-07-18_18-03-40.png?generation=1563469584580239&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Test</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2Fc42a3eeda53de8562249b57d8eca69da%2Frsz_screenshot_from_2019-07-18_18-05-27.png?generation=1563469610737268&amp;alt=media\" alt=\"\"></p>\n\n<p>Looks to me a large number of the test images have been pre-processed: you can see they are cropped as there is a little black in the corners. I've also noticed that applying the winning techniques from the last competition (which applies some cropping) tends to increase local cv while reducing test performance.</p>\n\n<p>I wonder what the real test dataset will look like?</p>",
      "rawMarkdown": "This is very strange. Using exactly the same cv fold splits each time, I've had the following results:\n\n|Model   |fold0|\tfold1|\tfold2|\tfold3|\tfold4|\tCV|\tLB|\n| --- | --- | --- | --- | --- | --- |\nModel 0\t|0.8970\t|0.8983|\t0.9024\t|0.8920\t|0.8874\t|0.8954\t|0.80|\nModel 1\t|0.9015|\t0.9052|\t0.9085|\t0.9049\t|0.8869|\t0.9014|\t0.77|\nModel 2\t|0.9251|\t0.9158|\t0.9288|\t0.9180|\t0.9168|\t0.9209\t|0.79|\n\nThe  training set has 3662 pictures, so each validation fold has approx. 730 images. The test data has 1928 images. Yet we're seeing way more unstable results on the test data than we are across folds. Something is very different about the test data (which people have already noted, for sure).\n\nEven when you eyeball it (😒 ) you can see a pronounced difference between train and test:\n\n**Train**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2F1d6b371c0112d928a782b52b2f265287%2Frsz_screenshot_from_2019-07-18_18-03-40.png?generation=1563469584580239&amp;alt=media)\n\n**Test**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2Fc42a3eeda53de8562249b57d8eca69da%2Frsz_screenshot_from_2019-07-18_18-05-27.png?generation=1563469610737268&amp;alt=media)\n\nLooks to me a large number of the test images have been pre-processed: you can see they are cropped as there is a little black in the corners. I've also noticed that applying the winning techniques from the last competition (which applies some cropping) tends to increase local cv while reducing test performance.\n\nI wonder what the real test dataset will look like?\n\n\n",
      "votes": 20,
      "replies": [
        {
          "id": 580159,
          "postDate": "2019-07-19T18:53:55.763Z",
          "content": "<p>Did you pretrain on the previous competition data?</p>",
          "rawMarkdown": "Did you pretrain on the previous competition data?"
        },
        {
          "id": 580724,
          "postDate": "2019-07-20T16:15:46.110Z",
          "content": "<p>Very interesting this hypothesis about the test data being preprocessed, maybe we can assume that the private test set is similar.</p>",
          "rawMarkdown": "Very interesting this hypothesis about the test data being preprocessed, maybe we can assume that the private test set is similar."
        },
        {
          "id": 602239,
          "postDate": "2019-08-18T20:00:56.360Z",
          "content": "<p>Very strange indeed! Thanks for sharing your research. I should have taken a better look at the test data, haha!</p>",
          "rawMarkdown": "Very strange indeed! Thanks for sharing your research. I should have taken a better look at the test data, haha!",
          "votes": 1
        },
        {
          "id": 604286,
          "postDate": "2019-08-21T08:13:11.440Z",
          "content": "<p>A side question: are you using Manjaro as your operating system ?  I think this is a screenshot from Dolphin...</p>",
          "rawMarkdown": "A side question: are you using Manjaro as your operating system ?  I think this is a screenshot from Dolphin..."
        }
      ]
    },
    {
      "id": 573782,
      "postDate": "2019-07-12T18:27:22.317Z",
      "content": "<p>I made a comprehensive table of all submissions same validation set across experiments (without classification) ... I don't see any correlation ....</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F5df09394601c5e81b2e845504ef93a7e%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562956138366610&amp;alt=media\" alt=\"\"></p>\n\n<p>also here is a plot\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F244090f0ae2d974c16ff84335b1ede78%2Fimage.png?generation=1562956028368873&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I made a comprehensive table of all submissions same validation set across experiments (without classification) ... I don't see any correlation ....\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F5df09394601c5e81b2e845504ef93a7e%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562956138366610&amp;alt=media)\n\n\nalso here is a plot\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F244090f0ae2d974c16ff84335b1ede78%2Fimage.png?generation=1562956028368873&amp;alt=media)\n",
      "votes": 15,
      "replies": [
        {
          "id": 573878,
          "postDate": "2019-07-12T22:09:10.103Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> </p>\n\n<p>I think Quadratic Weighted Kappa is an unstable evaluation metric when train and test are different.\nHere is an experimental QWK distribution on sampled evaluation data (size: 549 = 0.15 * len(train) and 1098). OOF QWK score of my naive model with 5-fold CV is 0.922(It's almost a median of the distribution). We can see less examples for evaluation(= bigger difference of train and evaluation data) shows the fluctuation of QWK score. <br>\nMy current hypothesis is that public test data consists of higher severity examples than those of train. So even if we improve the performance of models on lower severity examples, it will not push our public LB scores.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5176d8fd46a55f19e559872b97c96b67%2Fqwk_dist_015.png?generation=1562968719248609&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F44b395f71073627a985a427861ab49dc%2Fqwk_dist_03.png?generation=1562968734228367&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@drhabib \n\nI think Quadratic Weighted Kappa is an unstable evaluation metric when train and test are different.\nHere is an experimental QWK distribution on sampled evaluation data (size: 549 = 0.15 * len(train) and 1098). OOF QWK score of my naive model with 5-fold CV is 0.922(It's almost a median of the distribution). We can see less examples for evaluation(= bigger difference of train and evaluation data) shows the fluctuation of QWK score.  \nMy current hypothesis is that public test data consists of higher severity examples than those of train. So even if we improve the performance of models on lower severity examples, it will not push our public LB scores.  \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5176d8fd46a55f19e559872b97c96b67%2Fqwk_dist_015.png?generation=1562968719248609&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F44b395f71073627a985a427861ab49dc%2Fqwk_dist_03.png?generation=1562968734228367&amp;alt=media)\n",
          "votes": 10
        },
        {
          "id": 573896,
          "postDate": "2019-07-12T23:11:23.137Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Woow thank you for the nice explanation and your efforts in making the plots!  </p>",
          "rawMarkdown": "@maxwell110 Woow thank you for the nice explanation and your efforts in making the plots!  ",
          "votes": 1
        },
        {
          "id": 573948,
          "postDate": "2019-07-13T02:27:46.100Z",
          "content": "<p>Maybe it is better to validate based on another metric that generalizes better. Make sense?</p>",
          "rawMarkdown": "Maybe it is better to validate based on another metric that generalizes better. Make sense?",
          "votes": 2
        },
        {
          "id": 578157,
          "postDate": "2019-07-17T12:08:58.153Z",
          "content": "<p>Really important information.\nSo...  what is the better metric than QWK for making more reliable cv?</p>",
          "rawMarkdown": "Really important information.\nSo...  what is the better metric than QWK for making more reliable cv?"
        },
        {
          "id": 579817,
          "postDate": "2019-07-19T08:58:47.160Z",
          "content": "<p>Maybe since the performance of QWK looks roughly normal we can just do a t-test to see if our new LB score is significantly better than an old LB score - that way we can be like 95% confident that we have actually improved, likewise do this on the CV scores. This doesn't solve the validation - test mismatch but it would at least help to deal with the noisiness of QWK.</p>",
          "rawMarkdown": "Maybe since the performance of QWK looks roughly normal we can just do a t-test to see if our new LB score is significantly better than an old LB score - that way we can be like 95% confident that we have actually improved, likewise do this on the CV scores. This doesn't solve the validation - test mismatch but it would at least help to deal with the noisiness of QWK.",
          "votes": 3
        },
        {
          "id": 598974,
          "postDate": "2019-08-14T10:06:47.243Z",
          "content": "<p>Wow <a href=\"/maxwell110\">@maxwell110</a> this is great! Thank you for the explanation!</p>",
          "rawMarkdown": "Wow @maxwell110 this is great! Thank you for the explanation!",
          "votes": 1
        }
      ]
    },
    {
      "id": 569128,
      "postDate": "2019-07-06T05:54:15.587Z",
      "content": "<p>I have tried <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a> this dataset and trained a pretrained model there. Validation result in that competition held in 2015 is around 0.78. And in this competition, CV: 0.9189, LB: 0.782. Used 6 folds seresnext101 and regression, no TTA. Since the top teams in <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard\">https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard</a> can get higher than 0.84, so there is a long way to go.</p>",
      "rawMarkdown": "I have tried https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized this dataset and trained a pretrained model there. Validation result in that competition held in 2015 is around 0.78. And in this competition, CV: 0.9189, LB: 0.782. Used 6 folds seresnext101 and regression, no TTA. Since the top teams in https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard can get higher than 0.84, so there is a long way to go.",
      "votes": 16,
      "replies": [
        {
          "id": 570553,
          "postDate": "2019-07-08T12:51:23.333Z",
          "content": "<p>May I ask how many epochs have you been training on the  old data? I find my model's learning speed is very slow on old data:D</p>",
          "rawMarkdown": "May I ask how many epochs have you been training on the  old data? I find my model's learning speed is very slow on old data:D"
        },
        {
          "id": 570626,
          "postDate": "2019-07-08T14:57:11.607Z",
          "content": "<p>I only used the kaggle kernel to train, and 11 epochs to finish pretraining, I simply used train-test splitting there. </p>",
          "rawMarkdown": "I only used the kaggle kernel to train, and 11 epochs to finish pretraining, I simply used train-test splitting there. ",
          "votes": 2
        },
        {
          "id": 570631,
          "postDate": "2019-07-08T15:03:43.663Z",
          "content": "<p>how did you get old competiton data in kaggle kernels? </p>",
          "rawMarkdown": "how did you get old competiton data in kaggle kernels? "
        },
        {
          "id": 570647,
          "postDate": "2019-07-08T15:19:19.900Z",
          "content": "<p>I imported this dataset: <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a>, really thanks to the author. This weekend I just timed out with my pretraining kernels twice since I add something new, wasted 18 hours😯 </p>",
          "rawMarkdown": "I imported this dataset: https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized, really thanks to the author. This weekend I just timed out with my pretraining kernels twice since I add something new, wasted 18 hours😯 ",
          "votes": 1
        },
        {
          "id": 572923,
          "postDate": "2019-07-11T15:16:13.163Z",
          "content": "<p>Hey! thanks for sharing that. How's that dataset's images different from this competition's images? Was there any specific preprocessing you had to do?</p>",
          "rawMarkdown": "Hey! thanks for sharing that. How's that dataset's images different from this competition's images? Was there any specific preprocessing you had to do?"
        }
      ]
    },
    {
      "id": 581042,
      "postDate": "2019-07-21T09:49:25.913Z",
      "content": "<p>CV: 0.885\nLB: 0.821</p>\n\n<p>image size 256x256\nEfficientNet-B3, no ensemble, no TTA\ntraining as a regression model\nusing both old and current competition dataset</p>",
      "rawMarkdown": "CV: 0.885\nLB: 0.821\n\nimage size 256x256\nEfficientNet-B3, no ensemble, no TTA\ntraining as a regression model\nusing both old and current competition dataset",
      "votes": 13,
      "replies": [
        {
          "id": 581046,
          "postDate": "2019-07-21T09:55:35.400Z",
          "content": "<p>Did you do some special kind of CV?</p>",
          "rawMarkdown": "Did you do some special kind of CV?",
          "votes": 1
        },
        {
          "id": 581065,
          "postDate": "2019-07-21T10:57:51.540Z",
          "content": "<p>Thanks for the info. It is really weird that lower CV seems to lead to higher LB and that LB scores vary that strongly. Pretrain and finetune or fitted together? Special type of CV/early stopping? Optimized kappa?</p>",
          "rawMarkdown": "Thanks for the info. It is really weird that lower CV seems to lead to higher LB and that LB scores vary that strongly. Pretrain and finetune or fitted together? Special type of CV/early stopping? Optimized kappa?"
        },
        {
          "id": 581184,
          "postDate": "2019-07-21T15:03:32.990Z",
          "content": "<p><a href=\"/abhishek\">@abhishek</a> \nI used about 20% of the aptos data for validation. I made to make the val distribution a little bid smooth because there were too many 0 labeled data. However, even if the random seed is changed, the score changes greatly, so I don't know what is essentially effective.\n<a href=\"/philippsinger\">@philippsinger</a> \nBoth old and current data are used together. The objective is MSE. monitoring Optimized kappa (to save).</p>",
          "rawMarkdown": "@abhishek \nI used about 20% of the aptos data for validation. I made to make the val distribution a little bid smooth because there were too many 0 labeled data. However, even if the random seed is changed, the score changes greatly, so I don't know what is essentially effective.\n@philippsinger \nBoth old and current data are used together. The objective is MSE. monitoring Optimized kappa (to save).",
          "votes": 5
        },
        {
          "id": 581909,
          "postDate": "2019-07-22T14:21:27.340Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> what does 'used together' mean?. My CV around 97.5 -&gt;&gt;&gt; LB: 77.5. CV 92 -&gt;&gt;&gt; LB 81.0 :)))</p>",
          "rawMarkdown": "@octpath0302 what does 'used together' mean?. My CV around 97.5 -&gt;&gt;&gt; LB: 77.5. CV 92 -&gt;&gt;&gt; LB 81.0 :)))"
        },
        {
          "id": 582589,
          "postDate": "2019-07-23T10:59:01.527Z",
          "content": "<p><a href=\"/dathudeptrai\">@dathudeptrai</a> After I constructed train / val from labeled Aptos data, I added labeled diabetic data to the train.\nIn my case, CV scores are always around 0.88.</p>",
          "rawMarkdown": "@dathudeptrai After I constructed train / val from labeled Aptos data, I added labeled diabetic data to the train.\nIn my case, CV scores are always around 0.88.",
          "votes": 1
        },
        {
          "id": 584016,
          "postDate": "2019-07-25T10:06:43.520Z",
          "content": "<p>I wonder you just did resize image and thats it?</p>",
          "rawMarkdown": "I wonder you just did resize image and thats it?"
        },
        {
          "id": 607456,
          "postDate": "2019-08-25T09:26:10.453Z",
          "content": "<p>@oct what best  qk you able to achieve with just old data i find that mingling old data send model for a toss. so when i do it with only old data i see qk is not going beyond 0.55 </p>",
          "rawMarkdown": "@oct what best  qk you able to achieve with just old data i find that mingling old data send model for a toss. so when i do it with only old data i see qk is not going beyond 0.55 "
        }
      ]
    },
    {
      "id": 580809,
      "postDate": "2019-07-20T19:49:46.227Z",
      "content": "<p>IMG SIZE 256, Ben grahm's color cropped images\nRegression, TTA, Optimized kappa, Pretraining on old competition dataset, finetuning on the given dataset\nsingle fold validation score: ~ 0.90, LB: 0.812</p>",
      "rawMarkdown": "IMG SIZE 256, Ben grahm's color cropped images\nRegression, TTA, Optimized kappa, Pretraining on old competition dataset, finetuning on the given dataset\nsingle fold validation score: ~ 0.90, LB: 0.812",
      "votes": 7
    },
    {
      "id": 580711,
      "postDate": "2019-07-20T16:02:26.237Z",
      "content": "<p>Update:\n```\n-5 fold CV\n-IMG SIZE 224\n-no TTA\n-no Optimized Kappa</p>\n\n<p>fold 0 - 0.916463\nfold 1- 0.906227\nfold 2- 0.917213\nfold 3- 0.929887\nfold 4-0.933146</p>\n\n<p>Average =LB 0.81\n```</p>",
      "rawMarkdown": "Update:\n```\n-5 fold CV\n-IMG SIZE 224\n-no TTA\n-no Optimized Kappa\n\n\nfold 0 - 0.916463\nfold 1- 0.906227\nfold 2- 0.917213\nfold 3- 0.929887\nfold 4-0.933146\n\nAverage =LB 0.81\n```",
      "votes": 5,
      "replies": [
        {
          "id": 580726,
          "postDate": "2019-07-20T16:18:19.073Z",
          "content": "<p>thx... is it regression + external data</p>",
          "rawMarkdown": "thx... is it regression + external data",
          "votes": 1
        },
        {
          "id": 580728,
          "postDate": "2019-07-20T16:20:17.837Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes",
          "votes": 1
        },
        {
          "id": 580730,
          "postDate": "2019-07-20T16:24:55.907Z",
          "content": "<p>how much boost regression gives you over classification, and how much external data? thx </p>",
          "rawMarkdown": "how much boost regression gives you over classification, and how much external data? thx ",
          "votes": 1
        },
        {
          "id": 580732,
          "postDate": "2019-07-20T16:36:39.617Z",
          "content": "<p>It's a significant boost in terms of how training goes.. . I haven't yet done proper experiments with classification, but I am planing to do next week, will report back =) </p>\n\n<p>In some experiment I used all the external data in other only to balance classes. In both cases one could get around 0.805-0.811. The question is how it will generalize to the private set...</p>",
          "rawMarkdown": "It's a significant boost in terms of how training goes.. . I haven't yet done proper experiments with classification, but I am planing to do next week, will report back =) \n\nIn some experiment I used all the external data in other only to balance classes. In both cases one could get around 0.805-0.811. The question is how it will generalize to the private set...",
          "votes": 1
        },
        {
          "id": 580737,
          "postDate": "2019-07-20T16:43:47.673Z",
          "content": "<p>thx. the question was did you try without external data. I did only one experiment with external data and I got 0.02 boost (i expect it to be more)</p>",
          "rawMarkdown": "thx. the question was did you try without external data. I did only one experiment with external data and I got 0.02 boost (i expect it to be more)",
          "votes": 1
        },
        {
          "id": 580738,
          "postDate": "2019-07-20T16:48:23.517Z",
          "content": "<p>hmm I have tried for older experiments. But I will try my latest experimental setup without external data and report back. </p>",
          "rawMarkdown": "hmm I have tried for older experiments. But I will try my latest experimental setup without external data and report back. ",
          "votes": 1
        },
        {
          "id": 580752,
          "postDate": "2019-07-20T17:31:03.803Z",
          "content": "<p>Hi drhb, are you doing any pre processing of the images at all?</p>",
          "rawMarkdown": "Hi drhb, are you doing any pre processing of the images at all?",
          "votes": 1
        },
        {
          "id": 580757,
          "postDate": "2019-07-20T17:33:05.090Z",
          "content": "<p>Not really... But one thing that I found important at least in my experiments was image augmentations during training. </p>\n\n<p>Hope it helps</p>",
          "rawMarkdown": "Not really... But one thing that I found important at least in my experiments was image augmentations during training. \n\nHope it helps"
        },
        {
          "id": 586082,
          "postDate": "2019-07-28T15:03:16.640Z",
          "content": "<p>hi ,drhb, what your validation data,use only new data or both new and old data for validation? If use only new data, the validation data is balanced or the same distribution with this competition train dataset?</p>",
          "rawMarkdown": "hi ,drhb, what your validation data,use only new data or both new and old data for validation? If use only new data, the validation data is balanced or the same distribution with this competition train dataset?"
        }
      ]
    },
    {
      "id": 581075,
      "postDate": "2019-07-21T11:19:12.987Z",
      "content": "<p>I think we should be very careful with reported CV scores since there is a clear issue in training data where image meta features are highly predictive of the target. I can get 0.70+ local kappa just by looking at image size and pixel counts (see <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510\">https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510</a>).</p>\n\n<p>IMO it is likely that a lot of higher local CV scores are inflated because of this.</p>",
      "rawMarkdown": "I think we should be very careful with reported CV scores since there is a clear issue in training data where image meta features are highly predictive of the target. I can get 0.70+ local kappa just by looking at image size and pixel counts (see https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510).\n\nIMO it is likely that a lot of higher local CV scores are inflated because of this.",
      "votes": 6
    },
    {
      "id": 604701,
      "postDate": "2019-08-21T17:04:27Z",
      "content": "<p>CV: 0.916, LB 0.795. No ensemble, no TTA, no optimization. Still fighting to minimize the gap. </p>",
      "rawMarkdown": "CV: 0.916, LB 0.795. No ensemble, no TTA, no optimization. Still fighting to minimize the gap. ",
      "votes": 4
    },
    {
      "id": 572202,
      "postDate": "2019-07-10T15:58:31.200Z",
      "content": "<p><code>\nmodel: ResNet50\nimge_sz: 256\nCV:  0.927678723\npublic LB: 0.771\n</code>\nno TTA, no x-fold CV</p>",
      "rawMarkdown": "```\nmodel: ResNet50\nimge_sz: 256\nCV:  0.927678723\npublic LB: 0.771\n```\nno TTA, no x-fold CV",
      "votes": 3,
      "replies": [
        {
          "id": 573027,
          "postDate": "2019-07-11T17:23:43.583Z",
          "content": "<p>Do you use the data of the previous challenge?. I train 80% training dataset and use 20% as valid (No CV). My kappa score on valid is 0.87, accuracy around 0.6, public LB 0.779. When i use other model, i got 0.91 kappa score on same valid set but public LB just 0.74. What do u think?. Do i need CV ?. This is the first time I play kaggle</p>",
          "rawMarkdown": "Do you use the data of the previous challenge?. I train 80% training dataset and use 20% as valid (No CV). My kappa score on valid is 0.87, accuracy around 0.6, public LB 0.779. When i use other model, i got 0.91 kappa score on same valid set but public LB just 0.74. What do u think?. Do i need CV ?. This is the first time I play kaggle",
          "votes": 3
        },
        {
          "id": 573038,
          "postDate": "2019-07-11T17:48:42.957Z",
          "content": "<p>Good Job on Achieving excellent score! I use only some images to Balance out for the missing classes.</p>\n\n<p>If you look this thread everybody is getting different scores, it seems like even slight seed can have a huge effect on the public leaderboard. It's not a good strategy to trust public standing. Of course  combining results across many folds is a good strategy... But I might be wrong,  I am like you a novice and trying to learn on Kaggle =) </p>",
          "rawMarkdown": "Good Job on Achieving excellent score! I use only some images to Balance out for the missing classes.\n\nIf you look this thread everybody is getting different scores, it seems like even slight seed can have a huge effect on the public leaderboard. It's not a good strategy to trust public standing. Of course  combining results across many folds is a good strategy... But I might be wrong,  I am like you a novice and trying to learn on Kaggle =) ",
          "votes": 4
        },
        {
          "id": 602192,
          "postDate": "2019-08-18T18:39:56.397Z",
          "content": "<p><a href=\"/dathudeptrai\">@dathudeptrai</a> <a href=\"/drhabib\">@drhabib</a>  <a href=\"/jionie\">@jionie</a> <a href=\"/suicaokhoailang\">@suicaokhoailang</a> <a href=\"/taindow\">@taindow</a> <a href=\"/octpath0302\">@octpath0302</a> <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> @everybody </p>\n\n<p>regarding the topic <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">\"be careful what you train on ...\"</a> and the scores you're reporting, are those models taking care of / using info related to image size ratio, or black borders?</p>\n\n<p>after removing the black borders, mi score dropped dramatically, so in my case it must be trusting these info</p>",
          "rawMarkdown": "@dathudeptrai @drhabib  @jionie @suicaokhoailang @taindow @octpath0302 @rishabhiitbhu @everybody \n\nregarding the topic [\"be careful what you train on ...\"](https://www.kaggle.com/taindow/be-careful-what-you-train-on) and the scores you're reporting, are those models taking care of / using info related to image size ratio, or black borders?\n\nafter removing the black borders, mi score dropped dramatically, so in my case it must be trusting these info\n\n\n\n\n\n\n"
        },
        {
          "id": 602206,
          "postDate": "2019-08-18T19:01:09.390Z",
          "content": "<p><a href=\"/virilo\">@virilo</a> I removed the black part, basically <code>crop_image_from_gray</code> function from <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a></p>",
          "rawMarkdown": "@virilo I removed the black part, basically `crop_image_from_gray` function from https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping",
          "votes": 1
        }
      ]
    },
    {
      "id": 572140,
      "postDate": "2019-07-10T14:25:30.677Z",
      "content": "<p>holdout (using 20% as valid) CV 0.9292 LB 0.767. ResNet101 with TTA</p>",
      "rawMarkdown": "holdout (using 20% as valid) CV 0.9292 LB 0.767. ResNet101 with TTA",
      "votes": 3
    },
    {
      "id": 566094,
      "postDate": "2019-07-01T19:35:57.620Z",
      "content": "<p>One possible explanation is that there are duplicated patients with multiple images in the training set ... </p>",
      "rawMarkdown": "One possible explanation is that there are duplicated patients with multiple images in the training set ... ",
      "votes": 3
    },
    {
      "id": 582306,
      "postDate": "2019-07-23T03:31:15.653Z",
      "content": "<p>5 fold with TTA.\nLB: 0.806\nCV: 0.818</p>",
      "rawMarkdown": "5 fold with TTA.\nLB: 0.806\nCV: 0.818",
      "votes": 4,
      "replies": [
        {
          "id": 582433,
          "postDate": "2019-07-23T07:12:15.117Z",
          "content": "<p>Hi Khoi,</p>\n\n<p>Really interesting you have such close CV/LB. Are you doing much pre-processing of the images?</p>",
          "rawMarkdown": "Hi Khoi,\n\nReally interesting you have such close CV/LB. Are you doing much pre-processing of the images?",
          "votes": 1
        },
        {
          "id": 582541,
          "postDate": "2019-07-23T09:51:02.813Z",
          "content": "<p>No special pre-processing tricks yet. Strangely for me those cropping and normalizing methods didn't work out well.</p>",
          "rawMarkdown": "No special pre-processing tricks yet. Strangely for me those cropping and normalizing methods didn't work out well.",
          "votes": 1
        },
        {
          "id": 582559,
          "postDate": "2019-07-23T10:16:20.093Z",
          "content": "<p>Regression? Is your CV stable across experiments</p>",
          "rawMarkdown": "Regression? Is your CV stable across experiments"
        },
        {
          "id": 582632,
          "postDate": "2019-07-23T11:32:02.787Z",
          "content": "<p>You need to tell what kind of data the CV is based on. My guess is 2015+2019 combined.</p>",
          "rawMarkdown": "You need to tell what kind of data the CV is based on. My guess is 2015+2019 combined.",
          "votes": 1
        },
        {
          "id": 582635,
          "postDate": "2019-07-23T11:36:56.057Z",
          "content": "<p>Correct. I think the key in this competition is to get creative with the external data, and pray the private data is what one think it would look like. \nThe difference between the training and public test set was already ridiculous that I won't be too surprised if the private one will be full of grayscale images with a few of dog photos just to screw you up.</p>",
          "rawMarkdown": "Correct. I think the key in this competition is to get creative with the external data, and pray the private data is what one think it would look like. \nThe difference between the training and public test set was already ridiculous that I won't be too surprised if the private one will be full of grayscale images with a few of dog photos just to screw you up.",
          "votes": 19
        },
        {
          "id": 582670,
          "postDate": "2019-07-23T12:27:16.890Z",
          "content": "<p>You are so funny, <a href=\"/suicaokhoailang\">@suicaokhoailang</a> .\nBTW I'm now afraid that the private test dataset includes some duplicated images with each different label. I have asked <a href=\"/sohier\">@sohier</a> in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97606#570261\">another discussion thread</a> about this, but not yet gotten any response. <br>\nLet's pray our train data would be same as the private one.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fa67741959f2c5ba0d1d44318feaacdd7%2Fpray01.jpg?generation=1563884810862041&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "You are so funny, @suicaokhoailang .\nBTW I'm now afraid that the private test dataset includes some duplicated images with each different label. I have asked @sohier in [another discussion thread](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97606#570261) about this, but not yet gotten any response.   \nLet's pray our train data would be same as the private one.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fa67741959f2c5ba0d1d44318feaacdd7%2Fpray01.jpg?generation=1563884810862041&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 573017,
      "postDate": "2019-07-11T17:00:27.113Z",
      "content": "<p>model: Xception\nimage_sz: 224\n20% as valid: kappa score around 87\npublic LB: 77.9</p>\n\n<p>model: Xception with some modify\nimage_sz: 224\n20% as valid (same valid set above): kappa score around 91\npublic LB: 74.0</p>\n\n<p>I don't know which one to trust :))</p>",
      "rawMarkdown": "model: Xception\nimage_sz: 224\n20% as valid: kappa score around 87\npublic LB: 77.9\n\nmodel: Xception with some modify\nimage_sz: 224\n20% as valid (same valid set above): kappa score around 91\npublic LB: 74.0\n\nI don't know which one to trust :))\n",
      "votes": 4,
      "replies": [
        {
          "id": 575400,
          "postDate": "2019-07-15T11:33:10.573Z",
          "content": "<p>Thanks for sharing ! \nAre you treating this as regression or classification problem ? </p>",
          "rawMarkdown": "Thanks for sharing ! \nAre you treating this as regression or classification problem ? "
        }
      ]
    },
    {
      "id": 607354,
      "postDate": "2019-08-25T04:34:49.313Z",
      "content": "<p>1fold-CV 0.814, LB 0.799. Will try ensemble and more folds.</p>",
      "rawMarkdown": "1fold-CV 0.814, LB 0.799. Will try ensemble and more folds.",
      "votes": 1,
      "replies": [
        {
          "id": 607355,
          "postDate": "2019-08-25T04:42:32.667Z",
          "content": "<p>Beautifully  small gap between cv and lb!!\nDid you use old competition dataset?</p>",
          "rawMarkdown": "Beautifully  small gap between cv and lb!!\nDid you use old competition dataset?"
        },
        {
          "id": 607369,
          "postDate": "2019-08-25T05:10:17.333Z",
          "content": "<p>Yes, and my model is regression.</p>",
          "rawMarkdown": "Yes, and my model is regression.",
          "votes": 1
        },
        {
          "id": 607397,
          "postDate": "2019-08-25T06:55:16.537Z",
          "content": "<p>Thx! And... did you apply any special preprocessing to achieve this small gap?\nI'm still struggling to lower the gap, haha.\nWhat do you think mainly contributes this small gap?</p>",
          "rawMarkdown": "Thx! And... did you apply any special preprocessing to achieve this small gap?\nI'm still struggling to lower the gap, haha.\nWhat do you think mainly contributes this small gap?",
          "votes": 1
        },
        {
          "id": 607462,
          "postDate": "2019-08-25T09:54:54.500Z",
          "content": "<p><a href=\"/takehiro\">@takehiro</a> did you entire old data or subset of it  only ?\ni use some part of so that all levels have comparable no of examples but i find that cv score not going beyond .55 so far with just old data.</p>",
          "rawMarkdown": "@takehiro did you entire old data or subset of it  only ?\ni use some part of so that all levels have comparable no of examples but i find that cv score not going beyond .55 so far with just old data."
        }
      ]
    },
    {
      "id": 586258,
      "postDate": "2019-07-28T22:52:30.343Z",
      "content": "<p>Sorry, I am pretty new to this stuff. LB stands for leader board score right? So like the score kaggle tells you after you submit a kernel. Does CV stand for cross validation score? What if you dont do cross validation and are just submitting one model runs prediction? </p>",
      "rawMarkdown": "Sorry, I am pretty new to this stuff. LB stands for leader board score right? So like the score kaggle tells you after you submit a kernel. Does CV stand for cross validation score? What if you dont do cross validation and are just submitting one model runs prediction? ",
      "votes": 1
    },
    {
      "id": 574571,
      "postDate": "2019-07-14T05:22:24.133Z",
      "content": "<p>Reading these comments and checking my own model outputs, I think the reasons of the difference between CV and LB are:\n- the quadratic cohen kappa score is not stable, i.e., calculating in small batches and averaging give different result when calculate in one large batch.\n- the public test set distribution is way too different then given training set. Public test set might be composed of 40% severe NPDR, 20% middle NPDR...</p>",
      "rawMarkdown": "Reading these comments and checking my own model outputs, I think the reasons of the difference between CV and LB are:\n- the quadratic cohen kappa score is not stable, i.e., calculating in small batches and averaging give different result when calculate in one large batch.\n- the public test set distribution is way too different then given training set. Public test set might be composed of 40% severe NPDR, 20% middle NPDR...",
      "votes": 1,
      "replies": [
        {
          "id": 574854,
          "postDate": "2019-07-14T15:13:21Z",
          "content": "<p>Do you intuit the private dataset is more similar to the training set or to the public one?</p>",
          "rawMarkdown": "Do you intuit the private dataset is more similar to the training set or to the public one?"
        },
        {
          "id": 575034,
          "postDate": "2019-07-14T23:43:10.813Z",
          "content": "<p>The private training dataset might be more similar to the public one, but larger in size (X3 or X4) I guess.</p>",
          "rawMarkdown": "The private training dataset might be more similar to the public one, but larger in size (X3 or X4) I guess."
        },
        {
          "id": 575035,
          "postDate": "2019-07-14T23:45:41.203Z",
          "content": "<p>And private test set might be more similar to the public test test.😄 </p>",
          "rawMarkdown": "And private test set might be more similar to the public test test.😄 "
        }
      ]
    },
    {
      "id": 568238,
      "postDate": "2019-07-04T15:23:14.340Z",
      "content": "<p>do we know how many images are on private test data ? Also public leaderboard is calculated only based on 13%.... I am wondering how trustworthy is public standing ?</p>\n\n<p><code>\nmodel: ResNet50\nimge_sz: 224\nvalidation kappa: 0.92853\noptimized validation kappa: 0.93229102\npublic LB: 0.755\n</code></p>\n\n<p>no TTA</p>",
      "rawMarkdown": "do we know how many images are on private test data ? Also public leaderboard is calculated only based on 13%.... I am wondering how trustworthy is public standing ?\n\n```\nmodel: ResNet50\nimge_sz: 224\nvalidation kappa: 0.92853\noptimized validation kappa: 0.93229102\npublic LB: 0.755\n```\n\nno TTA\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 568530,
          "postDate": "2019-07-05T04:27:56.857Z",
          "content": "<p>Public leaderboard calculated by 15% of test data and it has 1928 images. <br>\n15%-&gt;1928 images <br>\n85%-&gt;about 11000 images <br>\nI guess stage2 test data has about 11000 images.  </p>",
          "rawMarkdown": "Public leaderboard calculated by 15% of test data and it has 1928 images.  \n15%-&gt;1928 images  \n85%-&gt;about 11000 images  \nI guess stage2 test data has about 11000 images.  ",
          "votes": 1
        },
        {
          "id": 568591,
          "postDate": "2019-07-05T07:07:28.187Z",
          "content": "<p>From the data description page:</p>\n\n<blockquote>\n  <p>You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately).</p>\n</blockquote>",
          "rawMarkdown": "From the data description page:\n\n&gt; You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately).",
          "votes": 2
        }
      ]
    },
    {
      "id": 567351,
      "postDate": "2019-07-03T11:40:49.870Z",
      "content": "<p>CV: 0.9169 LB:0.755\n5folds ResNet50</p>",
      "rawMarkdown": "CV: 0.9169 LB:0.755\n5folds ResNet50",
      "votes": 1,
      "replies": [
        {
          "id": 568561,
          "postDate": "2019-07-05T05:35:18.890Z",
          "content": "<p><a href=\"/takuok\">@takuok</a> how do you combine 5 folds? max voting or average?</p>",
          "rawMarkdown": "@takuok how do you combine 5 folds? max voting or average?"
        },
        {
          "id": 568568,
          "postDate": "2019-07-05T05:55:11.790Z",
          "content": "<p>Simple averaging.\nI used rank averaging on pet competition, so will try it.</p>",
          "rawMarkdown": "Simple averaging.\nI used rank averaging on pet competition, so will try it.",
          "votes": 2
        },
        {
          "id": 569080,
          "postDate": "2019-07-06T02:54:14.903Z",
          "content": "<p><a href=\"/takuok\">@takuok</a> Are you using regression or classification? </p>",
          "rawMarkdown": "@takuok Are you using regression or classification? ",
          "votes": 1
        },
        {
          "id": 569092,
          "postDate": "2019-07-06T03:31:19.707Z",
          "content": "<p>Regression with rmse loss.</p>",
          "rawMarkdown": "Regression with rmse loss."
        },
        {
          "id": 569779,
          "postDate": "2019-07-07T10:08:33.597Z",
          "content": "<p>what is the difference between simple averaging and rank averaging? It would much appreciated if you link me an example of using rank averaging :)</p>",
          "rawMarkdown": "what is the difference between simple averaging and rank averaging? It would much appreciated if you link me an example of using rank averaging :)"
        },
        {
          "id": 570166,
          "postDate": "2019-07-08T00:01:18.907Z",
          "content": "<p>Thanks for the information. I have a long way to go. <a href=\"/takuok\">@takuok</a> </p>",
          "rawMarkdown": "Thanks for the information. I have a long way to go. @takuok "
        },
        {
          "id": 570301,
          "postDate": "2019-07-08T05:29:39.990Z",
          "content": "<p><a href=\"https://www.kaggle.com/kosiew/rank-averaging-script\">This kernel</a> is an example of rank averaging. <br>\nOn this method, we treat target with ranking, then averaging it. <a href=\"/yangsaewon\">@yangsaewon</a> </p>",
          "rawMarkdown": "[This kernel](https://www.kaggle.com/kosiew/rank-averaging-script) is an example of rank averaging.  \nOn this method, we treat target with ranking, then averaging it. @yangsaewon ",
          "votes": 3
        },
        {
          "id": 570646,
          "postDate": "2019-07-08T15:18:59.417Z",
          "content": "<p>I love the kaggle spirit thank you so much!</p>",
          "rawMarkdown": "I love the kaggle spirit thank you so much!"
        }
      ]
    },
    {
      "id": 567161,
      "postDate": "2019-07-03T05:59:41.643Z",
      "content": "<p>My CV is 0.90 but public LB is 0.676. I am going to use TTA and hope it will increase public LB result. Intresting I am getting now quite the same results with ResNet 50 and ResNet 152 after 10-15 epochs of training.</p>",
      "rawMarkdown": "My CV is 0.90 but public LB is 0.676. I am going to use TTA and hope it will increase public LB result. Intresting I am getting now quite the same results with ResNet 50 and ResNet 152 after 10-15 epochs of training.",
      "votes": 1
    },
    {
      "id": 565617,
      "postDate": "2019-07-01T07:12:58.087Z",
      "content": "<p>My CV:0.912, LB:0.691, Not using TTA.</p>",
      "rawMarkdown": "My CV:0.912, LB:0.691, Not using TTA.",
      "votes": 1,
      "replies": [
        {
          "id": 565816,
          "postDate": "2019-07-01T12:09:34.987Z",
          "content": "<p><a href=\"/currypurin\">@currypurin</a> what image size did you use?</p>",
          "rawMarkdown": "@currypurin what image size did you use?"
        },
        {
          "id": 566305,
          "postDate": "2019-07-02T03:18:42.807Z",
          "content": "<p>I used your kernel code. thank you.</p>\n\n<p><code>\nimage = Image.open(img_name)\nimage = image.resize((256, 256), resample=Image.BILINEAR)\n</code></p>",
          "rawMarkdown": "I used your kernel code. thank you.\n\n```\nimage = Image.open(img_name)\nimage = image.resize((256, 256), resample=Image.BILINEAR)\n```"
        }
      ]
    },
    {
      "id": 571307,
      "postDate": "2019-07-09T13:35:54.777Z",
      "content": "<p>4-fold CV 0.9077 with std of 0.0045, LB is 0.625 for classification approach with TTA for Resnet18 backbone.</p>",
      "rawMarkdown": "4-fold CV 0.9077 with std of 0.0045, LB is 0.625 for classification approach with TTA for Resnet18 backbone.",
      "votes": 2
    },
    {
      "id": 568880,
      "postDate": "2019-07-05T15:52:33.660Z",
      "content": "<p>However, the distribution of the public test data does not necessarily represent the private test data. The private test data is much larger so they probably have different distributions. Therfore, analyzing test data does not help much, right?</p>",
      "rawMarkdown": "However, the distribution of the public test data does not necessarily represent the private test data. The private test data is much larger so they probably have different distributions. Therfore, analyzing test data does not help much, right?",
      "votes": 2,
      "replies": [
        {
          "id": 570656,
          "postDate": "2019-07-08T15:27:21.933Z",
          "content": "<p>Test data  analysis help you understand, that this competition without public leaderboard.</p>",
          "rawMarkdown": "Test data  analysis help you understand, that this competition without public leaderboard."
        }
      ]
    },
    {
      "id": 582855,
      "postDate": "2019-07-23T17:00:25.617Z",
      "content": "<p>Я не знакома с аргументом weights=\"quadratic\"в качестве метрики</p>",
      "rawMarkdown": "Я не знакома с аргументом weights=\"quadratic\"в качестве метрики",
      "votes": -7
    },
    {
      "id": 752253,
      "postDate": "2020-02-20T20:18:48.877Z",
      "content": "<p>good!</p>",
      "rawMarkdown": "good!"
    },
    {
      "id": 579446,
      "postDate": "2019-07-18T20:47:08.240Z",
      "content": "<p>0.91 cv unoptimized kappa vs 0.771 lb</p>",
      "rawMarkdown": "0.91 cv unoptimized kappa vs 0.771 lb"
    },
    {
      "id": 570398,
      "postDate": "2019-07-08T08:27:20.343Z",
      "content": "<p>Train with 0.9 dataset, valid with the rest 0.1. My CV is 0.91, but get 0.56 LB score. Terrible!</p>",
      "rawMarkdown": "Train with 0.9 dataset, valid with the rest 0.1. My CV is 0.91, but get 0.56 LB score. Terrible!"
    },
    {
      "id": 568875,
      "postDate": "2019-07-05T15:38:36Z",
      "content": "<p>Yeah, i've noticed that the greater my CV performance the worse my LB.  I've tried several different approaches, but my best is poor:</p>\n\n<p>CV:0.935 LB:0.704</p>\n\n<p>Given the poor confidence in the diagnosis itself, I am considering moving to an unsupervised method.</p>",
      "rawMarkdown": "Yeah, i've noticed that the greater my CV performance the worse my LB.  I've tried several different approaches, but my best is poor:\n\nCV:0.935 LB:0.704\n\nGiven the poor confidence in the diagnosis itself, I am considering moving to an unsupervised method.",
      "replies": [
        {
          "id": 569885,
          "postDate": "2019-07-07T13:54:58.013Z",
          "content": "<p>Try submit the lowest loss checkpoint.</p>",
          "rawMarkdown": "Try submit the lowest loss checkpoint.",
          "votes": 2
        }
      ]
    },
    {
      "id": 568859,
      "postDate": "2019-07-05T14:51:50.223Z",
      "content": "<p>So far, I am observing large differences between local CV and LB using resnet50:\n- 5-fold CV: 0.881\n- public LB: 0.615</p>\n\n<p>The gap seems to be higher compared to others. Interestingly, I also notice that my model tends to predict too few zeroes for the test data. </p>",
      "rawMarkdown": "So far, I am observing large differences between local CV and LB using resnet50:\n- 5-fold CV: 0.881\n- public LB: 0.615\n\nThe gap seems to be higher compared to others. Interestingly, I also notice that my model tends to predict too few zeroes for the test data. "
    },
    {
      "id": 566336,
      "postDate": "2019-07-02T04:26:08.810Z",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍 "
    },
    {
      "id": 566334,
      "postDate": "2019-07-02T04:25:48.293Z",
      "content": "<p>cool i here u bro</p>",
      "rawMarkdown": "cool i here u bro"
    },
    {
      "id": 565849,
      "postDate": "2019-07-01T13:03:31.270Z",
      "content": "<p>Considering the metric, it might be worth looking at the PetFinder Comp as well where we had the similar situation....</p>",
      "rawMarkdown": "Considering the metric, it might be worth looking at the PetFinder Comp as well where we had the similar situation...."
    },
    {
      "id": 564698,
      "postDate": "2019-06-29T20:44:34.997Z",
      "content": "<p>Well done! my CV is also 0.9 but lb is lower, 0.723. As for data distribution I totally agree; our training set is very small so next thing I'll try is data augmentation </p>",
      "rawMarkdown": "Well done! my CV is also 0.9 but lb is lower, 0.723. As for data distribution I totally agree; our training set is very small so next thing I'll try is data augmentation ",
      "replies": [
        {
          "id": 564702,
          "postDate": "2019-06-29T20:52:40.193Z",
          "content": "<p>TTA will improve your score. ive shown in this kernel how to do it: <a href=\"https://www.kaggle.com/abhishek/pytorch-inference-kernel\">https://www.kaggle.com/abhishek/pytorch-inference-kernel</a></p>",
          "rawMarkdown": "TTA will improve your score. ive shown in this kernel how to do it: https://www.kaggle.com/abhishek/pytorch-inference-kernel",
          "votes": 6
        }
      ]
    },
    {
      "id": 564547,
      "postDate": "2019-06-29T15:33:06.873Z",
      "content": "<p>bizarre phenomenon</p>",
      "rawMarkdown": "bizarre phenomenon"
    },
    {
      "id": 564430,
      "postDate": "2019-06-29T12:08:31.167Z",
      "content": "<p>just for the knowledge of beginners would you mind sharing ,what is cv score as lb is leaderboard</p>",
      "rawMarkdown": "just for the knowledge of beginners would you mind sharing ,what is cv score as lb is leaderboard",
      "replies": [
        {
          "id": 564446,
          "postDate": "2019-06-29T12:37:09.100Z",
          "content": "<p>Cross validation =CV,\nBtw i only trained single fold. Validation 0.9+ translated to LB 0.74+ w/o TTA.</p>\n\n<p>Update: my cv  is not stable 0.92+ &gt;&gt;&gt; from 0.68 to 0.758 on lb</p>",
          "rawMarkdown": "Cross validation =CV,\nBtw i only trained single fold. Validation 0.9+ translated to LB 0.74+ w/o TTA.\n\nUpdate: my cv  is not stable 0.92+ &gt;&gt;&gt; from 0.68 to 0.758 on lb",
          "votes": 4
        },
        {
          "id": 564787,
          "postDate": "2019-06-30T01:31:50.630Z",
          "content": "<p>Thanks <a href=\"/valanm\">@valanm</a> </p>",
          "rawMarkdown": "Thanks @valanm "
        },
        {
          "id": 565310,
          "postDate": "2019-06-30T18:41:17.130Z",
          "content": "<p>By CV \"score\" you mean accuracy or do you use a similar \"score\" function like kaggle ? (this is my first competition .. so I am new to this) </p>",
          "rawMarkdown": "By CV \"score\" you mean accuracy or do you use a similar \"score\" function like kaggle ? (this is my first competition .. so I am new to this) \n"
        },
        {
          "id": 565316,
          "postDate": "2019-06-30T18:52:25.647Z",
          "content": "<p>they mean the competition metric that is cohens kappa.</p>",
          "rawMarkdown": "they mean the competition metric that is cohens kappa."
        },
        {
          "id": 565333,
          "postDate": "2019-06-30T19:22:07.020Z",
          "content": "<p>Ok I have it in my metrics now ;-) only for printing I use best checkpoint of val_acc and the train is minimizing loss=cross_entropy.</p>\n\n<p>Should I use cohens kappa for something (as loss function or best_model criteria ..)  ? I just print - testing now and iteration 30 has value 0.55 ... ACC 0.77  I know it will reach acc 0.98 so cohens kappa will go up ... in val_acc I reach 0.75 ... (no cross-validation just a small random subset ).</p>",
          "rawMarkdown": "Ok I have it in my metrics now ;-) only for printing I use best checkpoint of val_acc and the train is minimizing loss=cross_entropy.\n\nShould I use cohens kappa for something (as loss function or best_model criteria ..)  ? I just print - testing now and iteration 30 has value 0.55 ... ACC 0.77  I know it will reach acc 0.98 so cohens kappa will go up ... in val_acc I reach 0.75 ... (no cross-validation just a small random subset ).\n"
        },
        {
          "id": 565338,
          "postDate": "2019-06-30T19:35:59Z",
          "content": "<p>You can use kappa score on validation for selecting your models. I don't know how it use it as loss. </p>",
          "rawMarkdown": "You can use kappa score on validation for selecting your models. I don't know how it use it as loss. "
        },
        {
          "id": 565349,
          "postDate": "2019-06-30T19:54:10.127Z",
          "content": "<p>maybe something like 1-cohens kappa for loss :-) if that has any logic at all... \nI have this in train now : \n 19s 7ms/step - loss: 0.0399 - acc: 0.9877 - cohens_kappa: 0.7645 - val_loss: 1.4088 - val_acc: 0.7463 - val_cohens_kappa: 0.7650</p>\n\n<p>It is a very simple model I just want to test If I can make a submission </p>",
          "rawMarkdown": "maybe something like 1-cohens kappa for loss :-) if that has any logic at all... \nI have this in train now : \n 19s 7ms/step - loss: 0.0399 - acc: 0.9877 - cohens_kappa: 0.7645 - val_loss: 1.4088 - val_acc: 0.7463 - val_cohens_kappa: 0.7650\n\nIt is a very simple model I just want to test If I can make a submission \n"
        },
        {
          "id": 565359,
          "postDate": "2019-06-30T20:14:48.947Z",
          "content": "<p>I am not sure (1-kappa) will work as loss function.  </p>",
          "rawMarkdown": "I am not sure (1-kappa) will work as loss function.  "
        },
        {
          "id": 566097,
          "postDate": "2019-07-01T19:38:15.563Z",
          "content": "<p>You cannot use kappa as loss but you do have soft versions of kappa which can estimate it.. (and be used as loss)</p>",
          "rawMarkdown": "You cannot use kappa as loss but you do have soft versions of kappa which can estimate it.. (and be used as loss)"
        },
        {
          "id": 566915,
          "postDate": "2019-07-02T19:15:52.803Z",
          "content": "<p>hello </p>\n\n<p>just to be sure, I am using this function for kappa metric: </p>\n\n<p><code>\ndef cohens_kappa(y_true, y_pred):\n    y_true_classes = tf.argmax(y_true, 1)\n    y_pred_classes = tf.argmax(y_pred, 1)\n    return tf.contrib.metrics.cohen_kappa(y_true_classes, y_pred_classes, 5)[1]\n</code></p>\n\n<p>this is correct right ? </p>\n\n<p>thanks</p>",
          "rawMarkdown": "hello \n\njust to be sure, I am using this function for kappa metric: \n\n```\ndef cohens_kappa(y_true, y_pred):\n    y_true_classes = tf.argmax(y_true, 1)\n    y_pred_classes = tf.argmax(y_pred, 1)\n    return tf.contrib.metrics.cohen_kappa(y_true_classes, y_pred_classes, 5)[1]\n```\n\nthis is correct right ? \n\nthanks"
        },
        {
          "id": 567207,
          "postDate": "2019-07-03T07:09:16.433Z",
          "content": "<p>I am not familiar with tensorflow. So can't comment !</p>",
          "rawMarkdown": "I am not familiar with tensorflow. So can't comment !"
        },
        {
          "id": 572942,
          "postDate": "2019-07-11T15:35:33.050Z",
          "content": "<p><a href=\"/birinhos\">@birinhos</a> use argument <code>weights=\"quadratic\"</code> as the metric for this competition is quadratic kappa score.</p>",
          "rawMarkdown": "@birinhos use argument `weights=\"quadratic\"` as the metric for this competition is quadratic kappa score.",
          "votes": 1
        }
      ]
    },
    {
      "id": 573947,
      "postDate": "2019-07-13T02:24:33.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 574045,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-07-13T06:42:09.147000",
      "content": "<p>5-fold CV: 0.8954\nLB: 80.0</p>\n\n<p>resnext101_32x16d /w pre-training on previous competition data and no TTA\ntreating as regression problem</p>",
      "votes": 22,
      "replies": [
        {
          "id": 575339,
          "author_name": "GSD",
          "author_url": "",
          "post_date": "2019-07-15T09:56:57.833000",
          "content": "<p>Hi Tom,just wanted to check how much time it took to train the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 575567,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-15T16:46:30.567000",
          "content": "<p>Hi GSD,</p>\n\n<p>img size = 256\nbatch size = 32\ngradient accumulation steps = 2\nmixed precision = True</p>\n\n<p>takes about 1.2 minute per epoch on RTX 2080ti (approx. 3000 images).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 575599,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-15T17:54:26.310000",
          "content": "<p>Did you finetune with imagenet weights or from scratch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 575633,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-15T18:41:09.090000",
          "content": "<p>I'm using the pre-trained instagram weights which can be found here: <a href=\"https://github.com/facebookresearch/WSL-Images\">https://github.com/facebookresearch/WSL-Images</a>.</p>\n\n<p>Very simple training procedure for the 0.80:</p>\n\n<p>1) Load pre-trained instagram weights for 32x16d model\n2) Fine tune on old competition data with early stopping based on current competition data\n3) Fine tune for small number of epoch on new competition data</p>\n\n<p>BTW congrats on your triple GM, I enjoyed the youtube video :)</p>",
          "votes": 23,
          "replies": []
        },
        {
          "id": 576726,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2019-07-15T23:55:13.517000",
          "content": "<p>Hi Tom, can you tell my about what \"early stopping based on current competition data\" means? I'm trying to pretrain with previous competition data. thanks for your sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 576772,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-07-16T01:57:00.767000",
          "content": "<p>Maybe while pretraining, use this new train set as validation set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 576800,
          "author_name": "GSD",
          "author_url": "",
          "post_date": "2019-07-16T03:13:52.520000",
          "content": "<p>Thanks Tom..Much helpful</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 576820,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2019-07-16T04:05:15.373000",
          "content": "<p><a href=\"/jionie\">@jionie</a> oh I get it. thanks for the reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 576898,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-07-16T06:42:21.027000",
          "content": "<p>Thanks Tom for ur insight!! U mentioned that u r treating it as regression problem...Is it normal(simple) Regression or Ordinal Regression?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 576913,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-16T06:59:12.610000",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  I just used simple regression with mse loss (same as Abhishek's kernel).\n<a href=\"/yangsaewon\">@yangsaewon</a> jionie is correct, i used the new data as a validation set for the old for the initial training.</p>\n\n<p>I've added a kernel using the pre-trained instagram weights if anyone is interested:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/instagram-to-aptos-resnext-101-32x8d\">https://www.kaggle.com/taindow/instagram-to-aptos-resnext-101-32x8d</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 576932,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-16T07:16:16.470000",
          "content": "<p>Are you guys doing early stopping with loss or kappa btw?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 577868,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-17T06:35:28.627000",
          "content": "<p>im early stopping by kappa but its not of any use it seems.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 577975,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2019-07-17T08:44:07.033000",
          "content": "<p>From my experience, early stopping with Kappa is quite unstable. Seems to me that using MSE works better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 578148,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-17T11:47:28.863000",
          "content": "<p>Yeah I've been using kappa for early stopping, but agree with others it seems relatively unstable.</p>\n\n<p>As a general update, increasingly local cv scores of &gt;0.90 have resulted in lower LB scores between 0.77 and 0.80. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 607459,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-25T09:41:02.633000",
          "content": "<p>@ tom while doing pre train using old data\n1) did you take all data or subset of it  so that ratio of all severity levels more or less similar\n2) how much kp score u got using old data on  Validation set</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 565061,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2019-06-30T11:26:22.727000",
      "content": "<p>I think the obvious difference of train and test data is the distribution of image size. <br>\nAnd more to say, the image size of the less risk retinopathy might be smaller.  </p>\n\n<p><strong>train image size</strong> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F6beafd10871da3a7b80595b874203218%2Ftrain_img_size.png?generation=1561893474326004&amp;alt=media\" alt=\"train image size\"></p>\n\n<p><strong>test image size</strong> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F64ae4c8e765fdba37d1dd25b9a3dff96%2Ftest_img_size.png?generation=1561893554727501&amp;alt=media\" alt=\"\"></p>",
      "votes": 21,
      "replies": [
        {
          "id": 565401,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-06-30T22:43:00.533000",
          "content": "<p>Agreed. I've looked at (actually looked at) a sample of the test set. And from that sample, it looks true. I plan on looking at some simple statistics and checking PSI later to get a better idea about what has changed exactly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565422,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-06-30T23:37:16.793000",
          "content": "<p>Yes.  </p>\n\n<p>&gt; And more to say, the image size of the less risk retinopathy might be smaller.  </p>\n\n<p>My above comment is from this pictures.  </p>\n\n<p><strong>Diagnosis vs Image Row Size</strong> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F0578f9965b6f809b6d5ae27da0f7f0f2%2Frow.png?generation=1561937448006275&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Diagnosis vs Image Column Size</strong> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5be6f48f77a2d77ea2ba1b0f4c4553d7%2Fcol.png?generation=1561937472495149&amp;alt=media\" alt=\"\"></p>\n\n<p>I think the way to resize and crop images would be the key to tackle this kind of problems ( TTA is also important ). <br>\nAnd more, 5th position solution in the previous competition had used the information of image size. But I'm not sure that solution was safe or not in the view of overfitting.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 565519,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-07-01T03:51:49.867000",
          "content": "<p>This is what the predicted distribution on the test set looks like for me:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2Fa86219e1c1eca9d598aaf69adb8bb66a%2FCapture2.PNG?generation=1561952985345841&amp;alt=media\" alt=\"\"></p>\n\n<p>It looks very different than the predicted values of the training set. So either I have lots of overfitting --which I probably do-- or the distributions are quite different.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 565524,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-01T04:27:17.953000",
          "content": "<p><a href=\"/puremath86\">@puremath86</a> \nThanks for sharing your inference. <br>\nIt looks you did not optimize your predicted values  with QWK. <br>\nIt is interesting for me to see the histogram after you applied optimization.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565526,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-07-01T04:43:54.423000",
          "content": "<p>I plan on treating the problem as a classification problem in the future --but for now I am treating it as a regression. Hopefully, the residual analysis will be enlightening. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565562,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-01T05:44:24.183000",
          "content": "<p>In <a href=\"https://www.kaggle.com/miklgr500/auto-encoder\">my research</a> with autoencoder I obtain the similar results, but not so clear for interpretation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 565564,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-07-01T05:50:10.513000",
          "content": "<p>Those are some interesting embeddings it learned. Two distinct groups of NR. I wonder why?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565568,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-01T06:00:23.100000",
          "content": "<p>All people have two eyes:)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 565574,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-07-01T06:11:32.537000",
          "content": "<p>Is that really the reason? That make's sense. And not all people have two eyes...<img src=\"https://i.pinimg.com/236x/57/d3/e7/57d3e7f7603d3f5adf2d64aca06df5ed.jpg\" alt=\"\"></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 565603,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-01T06:50:45.463000",
          "content": "<p>Lol :)\nSorry, I'm not formulate my idea very well.\nYes, the right and left eyes are different; they have different positions of the nerve and blood vessels.\nBut in the present state of my research, this is just a hypothesis that an autoencoder can separate two types of eyes. Maybe if I used the introduction of a loss of triplet, they would be the same, but then I influence the model with existing labels and can not make an independent conclusion with the distribution of labels about the data set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 566319,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2019-07-02T03:50:58.040000",
          "content": "<p>I don't know that it is helpful to be able to distinguish a left eye from a right eye --especially with all of the data augmentation you'll end up using. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 579277,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-07-18T17:09:09.960000",
      "content": "<p>This is very strange. Using exactly the same cv fold splits each time, I've had the following results:</p>\n\n<p>|Model   |fold0|    fold1|  fold2|  fold3|  fold4|  CV| LB|\n| --- | --- | --- | --- | --- | --- |\nModel 0 |0.8970 |0.8983|    0.9024  |0.8920 |0.8874 |0.8954 |0.80|\nModel 1 |0.9015|    0.9052| 0.9085| 0.9049  |0.8869|    0.9014| 0.77|\nModel 2 |0.9251|    0.9158| 0.9288| 0.9180| 0.9168| 0.9209  |0.79|</p>\n\n<p>The  training set has 3662 pictures, so each validation fold has approx. 730 images. The test data has 1928 images. Yet we're seeing way more unstable results on the test data than we are across folds. Something is very different about the test data (which people have already noted, for sure).</p>\n\n<p>Even when you eyeball it (😒 ) you can see a pronounced difference between train and test:</p>\n\n<p><strong>Train</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2F1d6b371c0112d928a782b52b2f265287%2Frsz_screenshot_from_2019-07-18_18-03-40.png?generation=1563469584580239&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Test</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2Fc42a3eeda53de8562249b57d8eca69da%2Frsz_screenshot_from_2019-07-18_18-05-27.png?generation=1563469610737268&amp;alt=media\" alt=\"\"></p>\n\n<p>Looks to me a large number of the test images have been pre-processed: you can see they are cropped as there is a little black in the corners. I've also noticed that applying the winning techniques from the last competition (which applies some cropping) tends to increase local cv while reducing test performance.</p>\n\n<p>I wonder what the real test dataset will look like?</p>",
      "votes": 20,
      "replies": [
        {
          "id": 580159,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-07-19T18:53:55.763000",
          "content": "<p>Did you pretrain on the previous competition data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 580724,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-07-20T16:15:46.110000",
          "content": "<p>Very interesting this hypothesis about the test data being preprocessed, maybe we can assume that the private test set is similar.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 602239,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-18T20:00:56.360000",
          "content": "<p>Very strange indeed! Thanks for sharing your research. I should have taken a better look at the test data, haha!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 604286,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2019-08-21T08:13:11.440000",
          "content": "<p>A side question: are you using Manjaro as your operating system ?  I think this is a screenshot from Dolphin...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 573782,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-07-12T18:27:22.317000",
      "content": "<p>I made a comprehensive table of all submissions same validation set across experiments (without classification) ... I don't see any correlation ....</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F5df09394601c5e81b2e845504ef93a7e%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562956138366610&amp;alt=media\" alt=\"\"></p>\n\n<p>also here is a plot\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F244090f0ae2d974c16ff84335b1ede78%2Fimage.png?generation=1562956028368873&amp;alt=media\" alt=\"\"></p>",
      "votes": 15,
      "replies": [
        {
          "id": 573878,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-12T22:09:10.103000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> </p>\n\n<p>I think Quadratic Weighted Kappa is an unstable evaluation metric when train and test are different.\nHere is an experimental QWK distribution on sampled evaluation data (size: 549 = 0.15 * len(train) and 1098). OOF QWK score of my naive model with 5-fold CV is 0.922(It's almost a median of the distribution). We can see less examples for evaluation(= bigger difference of train and evaluation data) shows the fluctuation of QWK score. <br>\nMy current hypothesis is that public test data consists of higher severity examples than those of train. So even if we improve the performance of models on lower severity examples, it will not push our public LB scores.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F5176d8fd46a55f19e559872b97c96b67%2Fqwk_dist_015.png?generation=1562968719248609&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F44b395f71073627a985a427861ab49dc%2Fqwk_dist_03.png?generation=1562968734228367&amp;alt=media\" alt=\"\"></p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 573896,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-12T23:11:23.137000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Woow thank you for the nice explanation and your efforts in making the plots!  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 573948,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-07-13T02:27:46.100000",
          "content": "<p>Maybe it is better to validate based on another metric that generalizes better. Make sense?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 578157,
          "author_name": "PGiN",
          "author_url": "",
          "post_date": "2019-07-17T12:08:58.153000",
          "content": "<p>Really important information.\nSo...  what is the better metric than QWK for making more reliable cv?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 579817,
          "author_name": "Hugo Dolan",
          "author_url": "",
          "post_date": "2019-07-19T08:58:47.160000",
          "content": "<p>Maybe since the performance of QWK looks roughly normal we can just do a t-test to see if our new LB score is significantly better than an old LB score - that way we can be like 95% confident that we have actually improved, likewise do this on the CV scores. This doesn't solve the validation - test mismatch but it would at least help to deal with the noisiness of QWK.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 598974,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-14T10:06:47.243000",
          "content": "<p>Wow <a href=\"/maxwell110\">@maxwell110</a> this is great! Thank you for the explanation!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 569128,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-07-06T05:54:15.587000",
      "content": "<p>I have tried <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a> this dataset and trained a pretrained model there. Validation result in that competition held in 2015 is around 0.78. And in this competition, CV: 0.9189, LB: 0.782. Used 6 folds seresnext101 and regression, no TTA. Since the top teams in <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard\">https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard</a> can get higher than 0.84, so there is a long way to go.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 570553,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-07-08T12:51:23.333000",
          "content": "<p>May I ask how many epochs have you been training on the  old data? I find my model's learning speed is very slow on old data:D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 570626,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-07-08T14:57:11.607000",
          "content": "<p>I only used the kaggle kernel to train, and 11 epochs to finish pretraining, I simply used train-test splitting there. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 570631,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-08T15:03:43.663000",
          "content": "<p>how did you get old competiton data in kaggle kernels? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 570647,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-07-08T15:19:19.900000",
          "content": "<p>I imported this dataset: <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a>, really thanks to the author. This weekend I just timed out with my pretraining kernels twice since I add something new, wasted 18 hours😯 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 572923,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-11T15:16:13.163000",
          "content": "<p>Hey! thanks for sharing that. How's that dataset's images different from this competition's images? Was there any specific preprocessing you had to do?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 581042,
      "author_name": "oct_path",
      "author_url": "",
      "post_date": "2019-07-21T09:49:25.913000",
      "content": "<p>CV: 0.885\nLB: 0.821</p>\n\n<p>image size 256x256\nEfficientNet-B3, no ensemble, no TTA\ntraining as a regression model\nusing both old and current competition dataset</p>",
      "votes": 13,
      "replies": [
        {
          "id": 581046,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-21T09:55:35.400000",
          "content": "<p>Did you do some special kind of CV?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581065,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-21T10:57:51.540000",
          "content": "<p>Thanks for the info. It is really weird that lower CV seems to lead to higher LB and that LB scores vary that strongly. Pretrain and finetune or fitted together? Special type of CV/early stopping? Optimized kappa?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581184,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-07-21T15:03:32.990000",
          "content": "<p><a href=\"/abhishek\">@abhishek</a> \nI used about 20% of the aptos data for validation. I made to make the val distribution a little bid smooth because there were too many 0 labeled data. However, even if the random seed is changed, the score changes greatly, so I don't know what is essentially effective.\n<a href=\"/philippsinger\">@philippsinger</a> \nBoth old and current data are used together. The objective is MSE. monitoring Optimized kappa (to save).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 581909,
          "author_name": "Nguyen Quan Anh Minh",
          "author_url": "",
          "post_date": "2019-07-22T14:21:27.340000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> what does 'used together' mean?. My CV around 97.5 -&gt;&gt;&gt; LB: 77.5. CV 92 -&gt;&gt;&gt; LB 81.0 :)))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 582589,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-07-23T10:59:01.527000",
          "content": "<p><a href=\"/dathudeptrai\">@dathudeptrai</a> After I constructed train / val from labeled Aptos data, I added labeled diabetic data to the train.\nIn my case, CV scores are always around 0.88.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 584016,
          "author_name": "cruigo",
          "author_url": "",
          "post_date": "2019-07-25T10:06:43.520000",
          "content": "<p>I wonder you just did resize image and thats it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607456,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-25T09:26:10.453000",
          "content": "<p>@oct what best  qk you able to achieve with just old data i find that mingling old data send model for a toss. so when i do it with only old data i see qk is not going beyond 0.55 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 580809,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2019-07-20T19:49:46.227000",
      "content": "<p>IMG SIZE 256, Ben grahm's color cropped images\nRegression, TTA, Optimized kappa, Pretraining on old competition dataset, finetuning on the given dataset\nsingle fold validation score: ~ 0.90, LB: 0.812</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 580711,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-07-20T16:02:26.237000",
      "content": "<p>Update:\n```\n-5 fold CV\n-IMG SIZE 224\n-no TTA\n-no Optimized Kappa</p>\n\n<p>fold 0 - 0.916463\nfold 1- 0.906227\nfold 2- 0.917213\nfold 3- 0.929887\nfold 4-0.933146</p>\n\n<p>Average =LB 0.81\n```</p>",
      "votes": 5,
      "replies": [
        {
          "id": 580726,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-20T16:18:19.073000",
          "content": "<p>thx... is it regression + external data</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580728,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:20:17.837000",
          "content": "<p>yes</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580730,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-20T16:24:55.907000",
          "content": "<p>how much boost regression gives you over classification, and how much external data? thx </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580732,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:36:39.617000",
          "content": "<p>It's a significant boost in terms of how training goes.. . I haven't yet done proper experiments with classification, but I am planing to do next week, will report back =) </p>\n\n<p>In some experiment I used all the external data in other only to balance classes. In both cases one could get around 0.805-0.811. The question is how it will generalize to the private set...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580737,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-20T16:43:47.673000",
          "content": "<p>thx. the question was did you try without external data. I did only one experiment with external data and I got 0.02 boost (i expect it to be more)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580738,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:48:23.517000",
          "content": "<p>hmm I have tried for older experiments. But I will try my latest experimental setup without external data and report back. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580752,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-20T17:31:03.803000",
          "content": "<p>Hi drhb, are you doing any pre processing of the images at all?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 580757,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T17:33:05.090000",
          "content": "<p>Not really... But one thing that I found important at least in my experiments was image augmentations during training. </p>\n\n<p>Hope it helps</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 586082,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-07-28T15:03:16.640000",
          "content": "<p>hi ,drhb, what your validation data,use only new data or both new and old data for validation? If use only new data, the validation data is balanced or the same distribution with this competition train dataset?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 581075,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-07-21T11:19:12.987000",
      "content": "<p>I think we should be very careful with reported CV scores since there is a clear issue in training data where image meta features are highly predictive of the target. I can get 0.70+ local kappa just by looking at image size and pixel counts (see <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510\">https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510</a>).</p>\n\n<p>IMO it is likely that a lot of higher local CV scores are inflated because of this.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 604701,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-08-21T17:04:27",
      "content": "<p>CV: 0.916, LB 0.795. No ensemble, no TTA, no optimization. Still fighting to minimize the gap. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 572202,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-07-10T15:58:31.200000",
      "content": "<p><code>\nmodel: ResNet50\nimge_sz: 256\nCV:  0.927678723\npublic LB: 0.771\n</code>\nno TTA, no x-fold CV</p>",
      "votes": 3,
      "replies": [
        {
          "id": 573027,
          "author_name": "Nguyen Quan Anh Minh",
          "author_url": "",
          "post_date": "2019-07-11T17:23:43.583000",
          "content": "<p>Do you use the data of the previous challenge?. I train 80% training dataset and use 20% as valid (No CV). My kappa score on valid is 0.87, accuracy around 0.6, public LB 0.779. When i use other model, i got 0.91 kappa score on same valid set but public LB just 0.74. What do u think?. Do i need CV ?. This is the first time I play kaggle</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 573038,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T17:48:42.957000",
          "content": "<p>Good Job on Achieving excellent score! I use only some images to Balance out for the missing classes.</p>\n\n<p>If you look this thread everybody is getting different scores, it seems like even slight seed can have a huge effect on the public leaderboard. It's not a good strategy to trust public standing. Of course  combining results across many folds is a good strategy... But I might be wrong,  I am like you a novice and trying to learn on Kaggle =) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 602192,
          "author_name": "Virilo Tejedor Aguilera",
          "author_url": "",
          "post_date": "2019-08-18T18:39:56.397000",
          "content": "<p><a href=\"/dathudeptrai\">@dathudeptrai</a> <a href=\"/drhabib\">@drhabib</a>  <a href=\"/jionie\">@jionie</a> <a href=\"/suicaokhoailang\">@suicaokhoailang</a> <a href=\"/taindow\">@taindow</a> <a href=\"/octpath0302\">@octpath0302</a> <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> @everybody </p>\n\n<p>regarding the topic <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">\"be careful what you train on ...\"</a> and the scores you're reporting, are those models taking care of / using info related to image size ratio, or black borders?</p>\n\n<p>after removing the black borders, mi score dropped dramatically, so in my case it must be trusting these info</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 602206,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-08-18T19:01:09.390000",
          "content": "<p><a href=\"/virilo\">@virilo</a> I removed the black part, basically <code>crop_image_from_gray</code> function from <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 572140,
      "author_name": "ynktk",
      "author_url": "",
      "post_date": "2019-07-10T14:25:30.677000",
      "content": "<p>holdout (using 20% as valid) CV 0.9292 LB 0.767. ResNet101 with TTA</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 566094,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "2019-07-01T19:35:57.620000",
      "content": "<p>One possible explanation is that there are duplicated patients with multiple images in the training set ... </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 582306,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2019-07-23T03:31:15.653000",
      "content": "<p>5 fold with TTA.\nLB: 0.806\nCV: 0.818</p>",
      "votes": 4,
      "replies": [
        {
          "id": 582433,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-23T07:12:15.117000",
          "content": "<p>Hi Khoi,</p>\n\n<p>Really interesting you have such close CV/LB. Are you doing much pre-processing of the images?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 582541,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-07-23T09:51:02.813000",
          "content": "<p>No special pre-processing tricks yet. Strangely for me those cropping and normalizing methods didn't work out well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 582559,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-23T10:16:20.093000",
          "content": "<p>Regression? Is your CV stable across experiments</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 582632,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-23T11:32:02.787000",
          "content": "<p>You need to tell what kind of data the CV is based on. My guess is 2015+2019 combined.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 582635,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-07-23T11:36:56.057000",
          "content": "<p>Correct. I think the key in this competition is to get creative with the external data, and pray the private data is what one think it would look like. \nThe difference between the training and public test set was already ridiculous that I won't be too surprised if the private one will be full of grayscale images with a few of dog photos just to screw you up.</p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 582670,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-23T12:27:16.890000",
          "content": "<p>You are so funny, <a href=\"/suicaokhoailang\">@suicaokhoailang</a> .\nBTW I'm now afraid that the private test dataset includes some duplicated images with each different label. I have asked <a href=\"/sohier\">@sohier</a> in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97606#570261\">another discussion thread</a> about this, but not yet gotten any response. <br>\nLet's pray our train data would be same as the private one.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fa67741959f2c5ba0d1d44318feaacdd7%2Fpray01.jpg?generation=1563884810862041&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 573017,
      "author_name": "Nguyen Quan Anh Minh",
      "author_url": "",
      "post_date": "2019-07-11T17:00:27.113000",
      "content": "<p>model: Xception\nimage_sz: 224\n20% as valid: kappa score around 87\npublic LB: 77.9</p>\n\n<p>model: Xception with some modify\nimage_sz: 224\n20% as valid (same valid set above): kappa score around 91\npublic LB: 74.0</p>\n\n<p>I don't know which one to trust :))</p>",
      "votes": 4,
      "replies": [
        {
          "id": 575400,
          "author_name": "Flemeille",
          "author_url": "",
          "post_date": "2019-07-15T11:33:10.573000",
          "content": "<p>Thanks for sharing ! \nAre you treating this as regression or classification problem ? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 607354,
      "author_name": "toto",
      "author_url": "",
      "post_date": "2019-08-25T04:34:49.313000",
      "content": "<p>1fold-CV 0.814, LB 0.799. Will try ensemble and more folds.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 607355,
          "author_name": "PGiN",
          "author_url": "",
          "post_date": "2019-08-25T04:42:32.667000",
          "content": "<p>Beautifully  small gap between cv and lb!!\nDid you use old competition dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607369,
          "author_name": "toto",
          "author_url": "",
          "post_date": "2019-08-25T05:10:17.333000",
          "content": "<p>Yes, and my model is regression.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 607397,
          "author_name": "PGiN",
          "author_url": "",
          "post_date": "2019-08-25T06:55:16.537000",
          "content": "<p>Thx! And... did you apply any special preprocessing to achieve this small gap?\nI'm still struggling to lower the gap, haha.\nWhat do you think mainly contributes this small gap?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 607462,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-25T09:54:54.500000",
          "content": "<p><a href=\"/takehiro\">@takehiro</a> did you entire old data or subset of it  only ?\ni use some part of so that all levels have comparable no of examples but i find that cv score not going beyond .55 so far with just old data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 586258,
      "author_name": "shatz",
      "author_url": "",
      "post_date": "2019-07-28T22:52:30.343000",
      "content": "<p>Sorry, I am pretty new to this stuff. LB stands for leader board score right? So like the score kaggle tells you after you submit a kernel. Does CV stand for cross validation score? What if you dont do cross validation and are just submitting one model runs prediction? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574571,
      "author_name": "V Zhou",
      "author_url": "",
      "post_date": "2019-07-14T05:22:24.133000",
      "content": "<p>Reading these comments and checking my own model outputs, I think the reasons of the difference between CV and LB are:\n- the quadratic cohen kappa score is not stable, i.e., calculating in small batches and averaging give different result when calculate in one large batch.\n- the public test set distribution is way too different then given training set. Public test set might be composed of 40% severe NPDR, 20% middle NPDR...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 574854,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-07-14T15:13:21",
          "content": "<p>Do you intuit the private dataset is more similar to the training set or to the public one?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 575034,
          "author_name": "V Zhou",
          "author_url": "",
          "post_date": "2019-07-14T23:43:10.813000",
          "content": "<p>The private training dataset might be more similar to the public one, but larger in size (X3 or X4) I guess.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 575035,
          "author_name": "V Zhou",
          "author_url": "",
          "post_date": "2019-07-14T23:45:41.203000",
          "content": "<p>And private test set might be more similar to the public test test.😄 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 568238,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-07-04T15:23:14.340000",
      "content": "<p>do we know how many images are on private test data ? Also public leaderboard is calculated only based on 13%.... I am wondering how trustworthy is public standing ?</p>\n\n<p><code>\nmodel: ResNet50\nimge_sz: 224\nvalidation kappa: 0.92853\noptimized validation kappa: 0.93229102\npublic LB: 0.755\n</code></p>\n\n<p>no TTA</p>",
      "votes": 1,
      "replies": [
        {
          "id": 568530,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2019-07-05T04:27:56.857000",
          "content": "<p>Public leaderboard calculated by 15% of test data and it has 1928 images. <br>\n15%-&gt;1928 images <br>\n85%-&gt;about 11000 images <br>\nI guess stage2 test data has about 11000 images.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 568591,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2019-07-05T07:07:28.187000",
          "content": "<p>From the data description page:</p>\n\n<blockquote>\n  <p>You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately).</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 567351,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2019-07-03T11:40:49.870000",
      "content": "<p>CV: 0.9169 LB:0.755\n5folds ResNet50</p>",
      "votes": 1,
      "replies": [
        {
          "id": 568561,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-05T05:35:18.890000",
          "content": "<p><a href=\"/takuok\">@takuok</a> how do you combine 5 folds? max voting or average?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 568568,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2019-07-05T05:55:11.790000",
          "content": "<p>Simple averaging.\nI used rank averaging on pet competition, so will try it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 569080,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-07-06T02:54:14.903000",
          "content": "<p><a href=\"/takuok\">@takuok</a> Are you using regression or classification? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 569092,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2019-07-06T03:31:19.707000",
          "content": "<p>Regression with rmse loss.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 569779,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2019-07-07T10:08:33.597000",
          "content": "<p>what is the difference between simple averaging and rank averaging? It would much appreciated if you link me an example of using rank averaging :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 570166,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-07-08T00:01:18.907000",
          "content": "<p>Thanks for the information. I have a long way to go. <a href=\"/takuok\">@takuok</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 570301,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2019-07-08T05:29:39.990000",
          "content": "<p><a href=\"https://www.kaggle.com/kosiew/rank-averaging-script\">This kernel</a> is an example of rank averaging. <br>\nOn this method, we treat target with ranking, then averaging it. <a href=\"/yangsaewon\">@yangsaewon</a> </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 570646,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2019-07-08T15:18:59.417000",
          "content": "<p>I love the kaggle spirit thank you so much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 567161,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2019-07-03T05:59:41.643000",
      "content": "<p>My CV is 0.90 but public LB is 0.676. I am going to use TTA and hope it will increase public LB result. Intresting I am getting now quite the same results with ResNet 50 and ResNet 152 after 10-15 epochs of training.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 565617,
      "author_name": "currypurin",
      "author_url": "",
      "post_date": "2019-07-01T07:12:58.087000",
      "content": "<p>My CV:0.912, LB:0.691, Not using TTA.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 565816,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-07-01T12:09:34.987000",
          "content": "<p><a href=\"/currypurin\">@currypurin</a> what image size did you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 566305,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2019-07-02T03:18:42.807000",
          "content": "<p>I used your kernel code. thank you.</p>\n\n<p><code>\nimage = Image.open(img_name)\nimage = image.resize((256, 256), resample=Image.BILINEAR)\n</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 571307,
      "author_name": "Eugene Khvedchenya",
      "author_url": "",
      "post_date": "2019-07-09T13:35:54.777000",
      "content": "<p>4-fold CV 0.9077 with std of 0.0045, LB is 0.625 for classification approach with TTA for Resnet18 backbone.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 568880,
      "author_name": "Carlos Prades K.",
      "author_url": "",
      "post_date": "2019-07-05T15:52:33.660000",
      "content": "<p>However, the distribution of the public test data does not necessarily represent the private test data. The private test data is much larger so they probably have different distributions. Therfore, analyzing test data does not help much, right?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 570656,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-08T15:27:21.933000",
          "content": "<p>Test data  analysis help you understand, that this competition without public leaderboard.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 582855,
      "author_name": "Marina",
      "author_url": "",
      "post_date": "2019-07-23T17:00:25.617000",
      "content": "<p>Я не знакома с аргументом weights=\"quadratic\"в качестве метрики</p>",
      "votes": -7,
      "replies": []
    },
    {
      "id": 752253,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T20:18:48.877000",
      "content": "<p>good!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579446,
      "author_name": "Nikolay Prokoptsev",
      "author_url": "",
      "post_date": "2019-07-18T20:47:08.240000",
      "content": "<p>0.91 cv unoptimized kappa vs 0.771 lb</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 570398,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-08T08:27:20.343000",
      "content": "<p>Train with 0.9 dataset, valid with the rest 0.1. My CV is 0.91, but get 0.56 LB score. Terrible!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 568875,
      "author_name": "Craig",
      "author_url": "",
      "post_date": "2019-07-05T15:38:36",
      "content": "<p>Yeah, i've noticed that the greater my CV performance the worse my LB.  I've tried several different approaches, but my best is poor:</p>\n\n<p>CV:0.935 LB:0.704</p>\n\n<p>Given the poor confidence in the diagnosis itself, I am considering moving to an unsupervised method.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 569885,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-07-07T13:54:58.013000",
          "content": "<p>Try submit the lowest loss checkpoint.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 568859,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-05T14:51:50.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 566336,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-02T04:26:08.810000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 566334,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-02T04:25:48.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 565849,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-01T13:03:31.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 564698,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-29T20:44:34.997000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 564702,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-29T20:52:40.193000",
          "content": "",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 564547,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-29T15:33:06.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 564430,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-29T12:08:31.167000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 564446,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-29T12:37:09.100000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 564787,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T01:31:50.630000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565310,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T18:41:17.130000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565316,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T18:52:25.647000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565333,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T19:22:07.020000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565338,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T19:35:59",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565349,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T19:54:10.127000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565359,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-30T20:14:48.947000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 566097,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-01T19:38:15.563000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 566915,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-02T19:15:52.803000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 567207,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-03T07:09:16.433000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 572942,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-11T15:35:33.050000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 573947,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-13T02:24:33.570000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "564326": "Let’s discuss CV and LB Scores here.\n\nTo start with, My CV: 0.90, LB: 0.745\n\nSo, it seems like the distribution of test data is not same as training data. What are your CV and LB scores ? ;)",
    "574045": "5-fold CV: 0.8954\nLB: 80.0\n\nresnext101_32x16d /w pre-training on previous competition data and no TTA\ntreating as regression problem",
    "565061": "I think the obvious difference of train and test data is the distribution of image size.  \nAnd more to say, the image size of the less risk retinopathy might be smaller.  \n\n**train image size**  \n![train image size](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F6beafd10871da3a7b80595b874203218%2Ftrain_img_size.png?generation=1561893474326004&amp;alt=media)\n\n**test image size**  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F64ae4c8e765fdba37d1dd25b9a3dff96%2Ftest_img_size.png?generation=1561893554727501&amp;alt=media)\n\n",
    "579277": "This is very strange. Using exactly the same cv fold splits each time, I've had the following results:\n\n|Model   |fold0|\tfold1|\tfold2|\tfold3|\tfold4|\tCV|\tLB|\n| --- | --- | --- | --- | --- | --- |\nModel 0\t|0.8970\t|0.8983|\t0.9024\t|0.8920\t|0.8874\t|0.8954\t|0.80|\nModel 1\t|0.9015|\t0.9052|\t0.9085|\t0.9049\t|0.8869|\t0.9014|\t0.77|\nModel 2\t|0.9251|\t0.9158|\t0.9288|\t0.9180|\t0.9168|\t0.9209\t|0.79|\n\nThe  training set has 3662 pictures, so each validation fold has approx. 730 images. The test data has 1928 images. Yet we're seeing way more unstable results on the test data than we are across folds. Something is very different about the test data (which people have already noted, for sure).\n\nEven when you eyeball it (😒 ) you can see a pronounced difference between train and test:\n\n**Train**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2F1d6b371c0112d928a782b52b2f265287%2Frsz_screenshot_from_2019-07-18_18-03-40.png?generation=1563469584580239&amp;alt=media)\n\n**Test**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F921947%2Fc42a3eeda53de8562249b57d8eca69da%2Frsz_screenshot_from_2019-07-18_18-05-27.png?generation=1563469610737268&amp;alt=media)\n\nLooks to me a large number of the test images have been pre-processed: you can see they are cropped as there is a little black in the corners. I've also noticed that applying the winning techniques from the last competition (which applies some cropping) tends to increase local cv while reducing test performance.\n\nI wonder what the real test dataset will look like?\n\n\n",
    "573782": "I made a comprehensive table of all submissions same validation set across experiments (without classification) ... I don't see any correlation ....\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F5df09394601c5e81b2e845504ef93a7e%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562956138366610&amp;alt=media)\n\n\nalso here is a plot\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F244090f0ae2d974c16ff84335b1ede78%2Fimage.png?generation=1562956028368873&amp;alt=media)\n",
    "569128": "I have tried https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized this dataset and trained a pretrained model there. Validation result in that competition held in 2015 is around 0.78. And in this competition, CV: 0.9189, LB: 0.782. Used 6 folds seresnext101 and regression, no TTA. Since the top teams in https://www.kaggle.com/c/diabetic-retinopathy-detection/leaderboard can get higher than 0.84, so there is a long way to go.",
    "581042": "CV: 0.885\nLB: 0.821\n\nimage size 256x256\nEfficientNet-B3, no ensemble, no TTA\ntraining as a regression model\nusing both old and current competition dataset",
    "580809": "IMG SIZE 256, Ben grahm's color cropped images\nRegression, TTA, Optimized kappa, Pretraining on old competition dataset, finetuning on the given dataset\nsingle fold validation score: ~ 0.90, LB: 0.812",
    "580711": "Update:\n```\n-5 fold CV\n-IMG SIZE 224\n-no TTA\n-no Optimized Kappa\n\n\nfold 0 - 0.916463\nfold 1- 0.906227\nfold 2- 0.917213\nfold 3- 0.929887\nfold 4-0.933146\n\nAverage =LB 0.81\n```",
    "581075": "I think we should be very careful with reported CV scores since there is a clear issue in training data where image meta features are highly predictive of the target. I can get 0.70+ local kappa just by looking at image size and pixel counts (see https://www.kaggle.com/taindow/be-careful-what-you-train-on?scriptVersionId=17552510).\n\nIMO it is likely that a lot of higher local CV scores are inflated because of this.",
    "604701": "CV: 0.916, LB 0.795. No ensemble, no TTA, no optimization. Still fighting to minimize the gap. ",
    "572202": "```\nmodel: ResNet50\nimge_sz: 256\nCV:  0.927678723\npublic LB: 0.771\n```\nno TTA, no x-fold CV",
    "572140": "holdout (using 20% as valid) CV 0.9292 LB 0.767. ResNet101 with TTA",
    "566094": "One possible explanation is that there are duplicated patients with multiple images in the training set ... ",
    "582306": "5 fold with TTA.\nLB: 0.806\nCV: 0.818",
    "573017": "model: Xception\nimage_sz: 224\n20% as valid: kappa score around 87\npublic LB: 77.9\n\nmodel: Xception with some modify\nimage_sz: 224\n20% as valid (same valid set above): kappa score around 91\npublic LB: 74.0\n\nI don't know which one to trust :))\n",
    "607354": "1fold-CV 0.814, LB 0.799. Will try ensemble and more folds.",
    "586258": "Sorry, I am pretty new to this stuff. LB stands for leader board score right? So like the score kaggle tells you after you submit a kernel. Does CV stand for cross validation score? What if you dont do cross validation and are just submitting one model runs prediction? ",
    "574571": "Reading these comments and checking my own model outputs, I think the reasons of the difference between CV and LB are:\n- the quadratic cohen kappa score is not stable, i.e., calculating in small batches and averaging give different result when calculate in one large batch.\n- the public test set distribution is way too different then given training set. Public test set might be composed of 40% severe NPDR, 20% middle NPDR...",
    "568238": "do we know how many images are on private test data ? Also public leaderboard is calculated only based on 13%.... I am wondering how trustworthy is public standing ?\n\n```\nmodel: ResNet50\nimge_sz: 224\nvalidation kappa: 0.92853\noptimized validation kappa: 0.93229102\npublic LB: 0.755\n```\n\nno TTA\n\n",
    "567351": "CV: 0.9169 LB:0.755\n5folds ResNet50",
    "567161": "My CV is 0.90 but public LB is 0.676. I am going to use TTA and hope it will increase public LB result. Intresting I am getting now quite the same results with ResNet 50 and ResNet 152 after 10-15 epochs of training.",
    "565617": "My CV:0.912, LB:0.691, Not using TTA.",
    "571307": "4-fold CV 0.9077 with std of 0.0045, LB is 0.625 for classification approach with TTA for Resnet18 backbone.",
    "568880": "However, the distribution of the public test data does not necessarily represent the private test data. The private test data is much larger so they probably have different distributions. Therfore, analyzing test data does not help much, right?",
    "582855": "Я не знакома с аргументом weights=\"quadratic\"в качестве метрики",
    "752253": "good!",
    "579446": "0.91 cv unoptimized kappa vs 0.771 lb",
    "570398": "Train with 0.9 dataset, valid with the rest 0.1. My CV is 0.91, but get 0.56 LB score. Terrible!",
    "568875": "Yeah, i've noticed that the greater my CV performance the worse my LB.  I've tried several different approaches, but my best is poor:\n\nCV:0.935 LB:0.704\n\nGiven the poor confidence in the diagnosis itself, I am considering moving to an unsupervised method.",
    "568859": "So far, I am observing large differences between local CV and LB using resnet50:\n- 5-fold CV: 0.881\n- public LB: 0.615\n\nThe gap seems to be higher compared to others. Interestingly, I also notice that my model tends to predict too few zeroes for the test data. ",
    "566336": "👍 ",
    "566334": "cool i here u bro",
    "565849": "Considering the metric, it might be worth looking at the PetFinder Comp as well where we had the similar situation....",
    "564698": "Well done! my CV is also 0.9 but lb is lower, 0.723. As for data distribution I totally agree; our training set is very small so next thing I'll try is data augmentation ",
    "564547": "bizarre phenomenon",
    "564430": "just for the knowledge of beginners would you mind sharing ,what is cv score as lb is leaderboard",
    "573947": ""
  }
}