{
  "id": 311887,
  "title": "How much does score gets better with bigger models and bigger image sizes?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/311887",
  "author_name": "Harshit Sheoran",
  "post_date": "2022-03-09T10:15:27.535000",
  "votes": 18,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Just like me I am sure, that other people here also are training their new approaches on a smaller test bench instead of a full fledged training to save massive time. How much score do  you gain after using bigger model or bigger image size?</p>\n<p>Scores below will be updated as I improve</p>\n<p>Let me start with mine:<br>\nB5 - 384 Single Fold (Usually the newest score, the later scores will be updated in a few days)<br>\nCV: 0.696<br>\nLB: 0.730</p>\n<p>B5 - 512 Single Fold<br>\nCV: 0.707<br>\nLB: 0.745</p>\n<p>B5 - 384 5 Folds<br>\nLB: 0.767</p>\n<p>B5 - 512 5 folds<br>\nLB: 0.778</p>\n<p>Scores above will be updated as I improve</p>",
  "messages": [
    {
      "id": 1716713,
      "postDate": "2022-03-09T10:15:27.537Z",
      "content": "<p>Just like me I am sure, that other people here also are training their new approaches on a smaller test bench instead of a full fledged training to save massive time. How much score do  you gain after using bigger model or bigger image size?</p>\n<p>Scores below will be updated as I improve</p>\n<p>Let me start with mine:<br>\nB5 - 384 Single Fold (Usually the newest score, the later scores will be updated in a few days)<br>\nCV: 0.696<br>\nLB: 0.730</p>\n<p>B5 - 512 Single Fold<br>\nCV: 0.707<br>\nLB: 0.745</p>\n<p>B5 - 384 5 Folds<br>\nLB: 0.767</p>\n<p>B5 - 512 5 folds<br>\nLB: 0.778</p>\n<p>Scores above will be updated as I improve</p>",
      "rawMarkdown": "Just like me I am sure, that other people here also are training their new approaches on a smaller test bench instead of a full fledged training to save massive time. How much score do  you gain after using bigger model or bigger image size?\n\nScores below will be updated as I improve\n\nLet me start with mine:\nB5 - 384 Single Fold (Usually the newest score, the later scores will be updated in a few days)\nCV: 0.696\nLB: 0.730\n\nB5 - 512 Single Fold\nCV: 0.707\nLB: 0.745\n\nB5 - 384 5 Folds\nLB: 0.767\n\nB5 - 512 5 folds\nLB: 0.778\n\nScores above will be updated as I improve",
      "votes": 18
    },
    {
      "id": 1716793,
      "postDate": "2022-03-09T12:04:17.030Z",
      "content": "<p>Yes, we can see correlation as well. <br>\nIn our case we manage to increse score about 0.02-0.03 jumping from 384 -&gt; 512. Certainly there is limitations we observe now:</p>\n<ul>\n<li>GPU/TPU memory -&gt; gradient exploding … to small batch size … to long training etc.</li>\n<li>increasing size (for small object) there is chance to loose some important image features (or change them significantly) - we managed to find golden point in our dataset (still looking for improvements) which gave us 0.810. </li>\n</ul>\n<p>I thnink that many of us use the same approach here (TF ArcFace/KNN or Pytorch ArcFace/GeM). The difference is mainly in two parts:</p>\n<ul>\n<li>dataset - we have different datasets (different ROI presented to NN). Moreover people from TOP3 (above 0.83) have better way to extract important features to NN.</li>\n<li>inference - some trick / postprocessing - I am still thinking about using some trick with feature matching - notebook I published - <a href=\"https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia\" target=\"_blank\">LoFTR notebook</a> (I tried SURF ans SIFT as well but they do not perform well in this situation).</li>\n</ul>",
      "rawMarkdown": "Yes, we can see correlation as well. \nIn our case we manage to increse score about 0.02-0.03 jumping from 384 -> 512. Certainly there is limitations we observe now:\n- GPU/TPU memory -> gradient exploding ... to small batch size ... to long training etc.\n- increasing size (for small object) there is chance to loose some important image features (or change them significantly) - we managed to find golden point in our dataset (still looking for improvements) which gave us 0.810. \n\nI thnink that many of us use the same approach here (TF ArcFace/KNN or Pytorch ArcFace/GeM). The difference is mainly in two parts:\n- dataset - we have different datasets (different ROI presented to NN). Moreover people from TOP3 (above 0.83) have better way to extract important features to NN.\n- inference - some trick / postprocessing - I am still thinking about using some trick with feature matching - notebook I published - [LoFTR notebook](https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia) (I tried SURF ans SIFT as well but they do not perform well in this situation).",
      "votes": 11,
      "replies": [
        {
          "id": 1716918,
          "postDate": "2022-03-09T14:06:47.763Z",
          "content": "<p>Thanks A Lot! for detailed explaination</p>\n<p>I saw gradient lose value, in my case, it certainly lowered the score a little but did not explode, I guess it did not explode because of accumulating gradients.</p>\n<p>I was also thinking about the same thing with using LoFTR features, have some plans to do on it but its time consuming as I dont have any previous experience in similar work.</p>\n<p>I dont think increasing size to a certain limit will cause any problem IF the image ROI is well enough.</p>\n<p>I personally use Pytorch ArcFace, and yes, although the dataset I use is completely public, a very specific way to mix it up gives me a boost of 0.02 points give or take.</p>\n<p>Inference, I am still improving score with this, personally I think this is where, LoFTR will provide benefits, another place to provide benefits with LoFTR is a very customised loss function, to make a good loss function, it would take a lot of behavioural information on the dataset.</p>",
          "rawMarkdown": "Thanks A Lot! for detailed explaination\n\nI saw gradient lose value, in my case, it certainly lowered the score a little but did not explode, I guess it did not explode because of accumulating gradients.\n\nI was also thinking about the same thing with using LoFTR features, have some plans to do on it but its time consuming as I dont have any previous experience in similar work.\n\nI dont think increasing size to a certain limit will cause any problem IF the image ROI is well enough.\n\nI personally use Pytorch ArcFace, and yes, although the dataset I use is completely public, a very specific way to mix it up gives me a boost of 0.02 points give or take.\n\nInference, I am still improving score with this, personally I think this is where, LoFTR will provide benefits, another place to provide benefits with LoFTR is a very customised loss function, to make a good loss function, it would take a lot of behavioural information on the dataset.",
          "votes": 2
        },
        {
          "id": 1718123,
          "postDate": "2022-03-10T14:19:28.520Z",
          "content": "<p>If you do not mind let me know if you use LoFTR or other feature matching techniques in your solution (it can be after competition certainly). </p>",
          "rawMarkdown": "If you do not mind let me know if you use LoFTR or other feature matching techniques in your solution (it can be after competition certainly). ",
          "votes": 1
        },
        {
          "id": 1718205,
          "postDate": "2022-03-10T15:41:31.990Z",
          "content": "<p>I will inform you if I end up using them, I am on a bottleneck right now, drastically improving score further is not to be seen currently, unless I make a really good dataset myself… which is likely my next step</p>",
          "rawMarkdown": "I will inform you if I end up using them, I am on a bottleneck right now, drastically improving score further is not to be seen currently, unless I make a really good dataset myself... which is likely my next step"
        },
        {
          "id": 1720924,
          "postDate": "2022-03-13T09:14:17.640Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1717540,
      "postDate": "2022-03-10T02:54:32.747Z",
      "content": "<p>model - &gt; swin_base<br>\nimage size -&gt; 224<br>\nLB = 0.512</p>",
      "rawMarkdown": "model - > swin_base\nimage size -> 224\nLB = 0.512\n",
      "votes": 1,
      "replies": [
        {
          "id": 1717561,
          "postDate": "2022-03-10T03:21:19.817Z",
          "content": "<p>Its really good to see that someone is able to use swin transformers, are you using tensorflow or pytorch?</p>",
          "rawMarkdown": "Its really good to see that someone is able to use swin transformers, are you using tensorflow or pytorch?"
        },
        {
          "id": 1717593,
          "postDate": "2022-03-10T04:13:29.423Z",
          "content": "<p>Using pytorch</p>",
          "rawMarkdown": "Using pytorch",
          "votes": 1
        }
      ]
    },
    {
      "id": 1718166,
      "postDate": "2022-03-10T15:03:48.863Z",
      "content": "<p>batch_size=?</p>",
      "rawMarkdown": "batch_size=?",
      "replies": [
        {
          "id": 1718203,
          "postDate": "2022-03-10T15:39:52.230Z",
          "content": "<p>accumulated batch size is 32 per gpu, gradient accumulation was used per 2 step with base batch size of 16 per gpu</p>",
          "rawMarkdown": "accumulated batch size is 32 per gpu, gradient accumulation was used per 2 step with base batch size of 16 per gpu",
          "votes": 1
        }
      ]
    },
    {
      "id": 1717903,
      "postDate": "2022-03-10T10:16:13.890Z",
      "content": "<p>Your result is single fold or ensemble of kfold?</p>",
      "rawMarkdown": "Your result is single fold or ensemble of kfold?",
      "replies": [
        {
          "id": 1717923,
          "postDate": "2022-03-10T10:47:45.670Z",
          "content": "<p>Single Fold</p>",
          "rawMarkdown": "Single Fold",
          "votes": 1
        }
      ]
    },
    {
      "id": 1717493,
      "postDate": "2022-03-10T01:18:42.200Z",
      "content": "<p>I've been using small models and img sizes to test out frameworks. </p>\n<p>B5 - 512<br>\nDataset: Detic Crops<br>\nLB: 0.695</p>\n<p>I suspect about a 0.04-0.07 increase from B5-B7 and larger image size.</p>",
      "rawMarkdown": "I've been using small models and img sizes to test out frameworks. \n\nB5 - 512\nDataset: Detic Crops\nLB: 0.695\n\nI suspect about a 0.04-0.07 increase from B5-B7 and larger image size.",
      "replies": [
        {
          "id": 1717508,
          "postDate": "2022-03-10T01:47:24.810Z",
          "content": "<p>Is this single fold or all folds?</p>",
          "rawMarkdown": "Is this single fold or all folds?"
        },
        {
          "id": 1717517,
          "postDate": "2022-03-10T01:59:20.720Z",
          "content": "<p>All folds. Is yours all folds?</p>",
          "rawMarkdown": "All folds. Is yours all folds?"
        },
        {
          "id": 1717524,
          "postDate": "2022-03-10T02:23:24.933Z",
          "content": "<p>Mine is single fold result, if yours were a single fold, I would asked to team with you, its still a really good score with detic crops (I am assuming its the same detic crops publicly open by phalanx) </p>",
          "rawMarkdown": "Mine is single fold result, if yours were a single fold, I would asked to team with you, its still a really good score with detic crops (I am assuming its the same detic crops publicly open by phalanx) "
        },
        {
          "id": 1717530,
          "postDate": "2022-03-10T02:39:13.233Z",
          "content": "<p>Nice. Still trying to figure out how to boost single fold score more but I have recently been trying out some new models and new methods. Are you using tensorflow or pytorch?</p>",
          "rawMarkdown": "Nice. Still trying to figure out how to boost single fold score more but I have recently been trying out some new models and new methods. Are you using tensorflow or pytorch?"
        },
        {
          "id": 1717535,
          "postDate": "2022-03-10T02:42:55.467Z",
          "content": "<p>Pytorch, I know it lowers the score a little in comparison to tensorflow (but idk why), but gives me way more freedom to implement whatever I feel like.</p>",
          "rawMarkdown": "Pytorch, I know it lowers the score a little in comparison to tensorflow (but idk why), but gives me way more freedom to implement whatever I feel like.",
          "votes": 1
        },
        {
          "id": 1717536,
          "postDate": "2022-03-10T02:47:20.303Z",
          "content": "<p>Oh ok. I've been using tensorflow due to performance difference in terms of speed on kaggle TPUs.</p>",
          "rawMarkdown": "Oh ok. I've been using tensorflow due to performance difference in terms of speed on kaggle TPUs."
        },
        {
          "id": 1718343,
          "postDate": "2022-03-10T18:31:00.907Z",
          "content": "<p>If you don't mind me asking what dataset have you used to get your score? </p>",
          "rawMarkdown": "If you don't mind me asking what dataset have you used to get your score? "
        },
        {
          "id": 1718593,
          "postDate": "2022-03-11T01:56:55.613Z",
          "content": "<p>Its a mix of datasets…, all publicaly available</p>",
          "rawMarkdown": "Its a mix of datasets..., all publicaly available",
          "votes": 2
        },
        {
          "id": 1718595,
          "postDate": "2022-03-11T01:58:57.780Z",
          "content": "<p>Thanks.      </p>",
          "rawMarkdown": "Thanks.      "
        },
        {
          "id": 1718599,
          "postDate": "2022-03-11T02:15:10.333Z",
          "content": "<p>Whats your single fold result, if you don't mind telling?</p>",
          "rawMarkdown": "Whats your single fold result, if you don't mind telling?"
        },
        {
          "id": 1718638,
          "postDate": "2022-03-11T03:29:50.413Z",
          "content": "<p>My single fold result is 0.649 on lb. Also my all fold result isn't what you think it is. ;)</p>",
          "rawMarkdown": "My single fold result is 0.649 on lb. Also my all fold result isn't what you think it is. ;)"
        },
        {
          "id": 1718683,
          "postDate": "2022-03-11T04:34:21.160Z",
          "content": "<p>There are too many possibilities to apply with \"Also my all fold result isn't what you think it is ;)\" I have posted a comment in looking for a team, if you match the requirenment, do reply there.</p>",
          "rawMarkdown": "There are too many possibilities to apply with \"Also my all fold result isn't what you think it is ;)\" I have posted a comment in looking for a team, if you match the requirenment, do reply there."
        },
        {
          "id": 1721044,
          "postDate": "2022-03-13T11:14:57.513Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> how much time is it taking per epoch in pytorch to train these models of yours</p>",
          "rawMarkdown": "@harshitsheoran how much time is it taking per epoch in pytorch to train these models of yours"
        },
        {
          "id": 1721050,
          "postDate": "2022-03-13T11:22:26.343Z",
          "content": "<p>I train on pytorch, my hardware is about 50-60% as good as tensorflow tpu, training 1 epoch takes just less than 3 minutes for image size 384, b5 model, I train it for 24 epochs, current result is 0.699 single fold</p>",
          "rawMarkdown": "I train on pytorch, my hardware is about 50-60% as good as tensorflow tpu, training 1 epoch takes just less than 3 minutes for image size 384, b5 model, I train it for 24 epochs, current result is 0.699 single fold",
          "votes": 1
        },
        {
          "id": 1732396,
          "postDate": "2022-03-23T10:57:02.653Z",
          "content": "<p>How much do you increase score with effnetb5 to b7? I only see an increase of 0.029 with the same hyperparameters.</p>",
          "rawMarkdown": "How much do you increase score with effnetb5 to b7? I only see an increase of 0.029 with the same hyperparameters."
        },
        {
          "id": 1732480,
          "postDate": "2022-03-23T12:30:50.450Z",
          "content": "<p>I experience less boost to increase score with bigger models (my boost is about 0.025?), as your score increases, its likely that the smaller models will catch up.</p>",
          "rawMarkdown": "I experience less boost to increase score with bigger models (my boost is about 0.025?), as your score increases, its likely that the smaller models will catch up."
        }
      ]
    },
    {
      "id": 1717205,
      "postDate": "2022-03-09T18:08:30.167Z",
      "content": "<p>Could someone provide insights on the gain we can expect by climbing the ladder from effnet-b0 to effnet-b7 ?</p>",
      "rawMarkdown": "Could someone provide insights on the gain we can expect by climbing the ladder from effnet-b0 to effnet-b7 ?\n\n "
    },
    {
      "id": 1716785,
      "postDate": "2022-03-09T11:56:19.457Z",
      "content": "<p>small model/image provide framework pipeline experiments, but definitely not good enough.</p>",
      "rawMarkdown": "small model/image provide framework pipeline experiments, but definitely not good enough."
    }
  ],
  "comments": [
    {
      "id": 1716793,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-03-09T12:04:17.030000",
      "content": "<p>Yes, we can see correlation as well. <br>\nIn our case we manage to increse score about 0.02-0.03 jumping from 384 -&gt; 512. Certainly there is limitations we observe now:</p>\n<ul>\n<li>GPU/TPU memory -&gt; gradient exploding … to small batch size … to long training etc.</li>\n<li>increasing size (for small object) there is chance to loose some important image features (or change them significantly) - we managed to find golden point in our dataset (still looking for improvements) which gave us 0.810. </li>\n</ul>\n<p>I thnink that many of us use the same approach here (TF ArcFace/KNN or Pytorch ArcFace/GeM). The difference is mainly in two parts:</p>\n<ul>\n<li>dataset - we have different datasets (different ROI presented to NN). Moreover people from TOP3 (above 0.83) have better way to extract important features to NN.</li>\n<li>inference - some trick / postprocessing - I am still thinking about using some trick with feature matching - notebook I published - <a href=\"https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia\" target=\"_blank\">LoFTR notebook</a> (I tried SURF ans SIFT as well but they do not perform well in this situation).</li>\n</ul>",
      "votes": 11,
      "replies": [
        {
          "id": 1716918,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-09T14:06:47.763000",
          "content": "<p>Thanks A Lot! for detailed explaination</p>\n<p>I saw gradient lose value, in my case, it certainly lowered the score a little but did not explode, I guess it did not explode because of accumulating gradients.</p>\n<p>I was also thinking about the same thing with using LoFTR features, have some plans to do on it but its time consuming as I dont have any previous experience in similar work.</p>\n<p>I dont think increasing size to a certain limit will cause any problem IF the image ROI is well enough.</p>\n<p>I personally use Pytorch ArcFace, and yes, although the dataset I use is completely public, a very specific way to mix it up gives me a boost of 0.02 points give or take.</p>\n<p>Inference, I am still improving score with this, personally I think this is where, LoFTR will provide benefits, another place to provide benefits with LoFTR is a very customised loss function, to make a good loss function, it would take a lot of behavioural information on the dataset.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1718123,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-03-10T14:19:28.520000",
          "content": "<p>If you do not mind let me know if you use LoFTR or other feature matching techniques in your solution (it can be after competition certainly). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1718205,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T15:41:31.990000",
          "content": "<p>I will inform you if I end up using them, I am on a bottleneck right now, drastically improving score further is not to be seen currently, unless I make a really good dataset myself… which is likely my next step</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1720924,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-03-13T09:14:17.640000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1717540,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-03-10T02:54:32.747000",
      "content": "<p>model - &gt; swin_base<br>\nimage size -&gt; 224<br>\nLB = 0.512</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1717561,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T03:21:19.817000",
          "content": "<p>Its really good to see that someone is able to use swin transformers, are you using tensorflow or pytorch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717593,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-03-10T04:13:29.423000",
          "content": "<p>Using pytorch</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1718166,
      "author_name": "AGEAGE",
      "author_url": "",
      "post_date": "2022-03-10T15:03:48.863000",
      "content": "<p>batch_size=?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1718203,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T15:39:52.230000",
          "content": "<p>accumulated batch size is 32 per gpu, gradient accumulation was used per 2 step with base batch size of 16 per gpu</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1717903,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2022-03-10T10:16:13.890000",
      "content": "<p>Your result is single fold or ensemble of kfold?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1717923,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T10:47:45.670000",
          "content": "<p>Single Fold</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1717493,
      "author_name": "Ari",
      "author_url": "",
      "post_date": "2022-03-10T01:18:42.200000",
      "content": "<p>I've been using small models and img sizes to test out frameworks. </p>\n<p>B5 - 512<br>\nDataset: Detic Crops<br>\nLB: 0.695</p>\n<p>I suspect about a 0.04-0.07 increase from B5-B7 and larger image size.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1717508,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T01:47:24.810000",
          "content": "<p>Is this single fold or all folds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717517,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-10T01:59:20.720000",
          "content": "<p>All folds. Is yours all folds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717524,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T02:23:24.933000",
          "content": "<p>Mine is single fold result, if yours were a single fold, I would asked to team with you, its still a really good score with detic crops (I am assuming its the same detic crops publicly open by phalanx) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717530,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-10T02:39:13.233000",
          "content": "<p>Nice. Still trying to figure out how to boost single fold score more but I have recently been trying out some new models and new methods. Are you using tensorflow or pytorch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717535,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-10T02:42:55.467000",
          "content": "<p>Pytorch, I know it lowers the score a little in comparison to tensorflow (but idk why), but gives me way more freedom to implement whatever I feel like.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1717536,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-10T02:47:20.303000",
          "content": "<p>Oh ok. I've been using tensorflow due to performance difference in terms of speed on kaggle TPUs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1718343,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-10T18:31:00.907000",
          "content": "<p>If you don't mind me asking what dataset have you used to get your score? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1718593,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-11T01:56:55.613000",
          "content": "<p>Its a mix of datasets…, all publicaly available</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1718595,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-11T01:58:57.780000",
          "content": "<p>Thanks.      </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1718599,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-11T02:15:10.333000",
          "content": "<p>Whats your single fold result, if you don't mind telling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1718638,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-11T03:29:50.413000",
          "content": "<p>My single fold result is 0.649 on lb. Also my all fold result isn't what you think it is. ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1718683,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-11T04:34:21.160000",
          "content": "<p>There are too many possibilities to apply with \"Also my all fold result isn't what you think it is ;)\" I have posted a comment in looking for a team, if you match the requirenment, do reply there.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1721044,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-03-13T11:14:57.513000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> how much time is it taking per epoch in pytorch to train these models of yours</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1721050,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-13T11:22:26.343000",
          "content": "<p>I train on pytorch, my hardware is about 50-60% as good as tensorflow tpu, training 1 epoch takes just less than 3 minutes for image size 384, b5 model, I train it for 24 epochs, current result is 0.699 single fold</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1732396,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2022-03-23T10:57:02.653000",
          "content": "<p>How much do you increase score with effnetb5 to b7? I only see an increase of 0.029 with the same hyperparameters.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1732480,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-23T12:30:50.450000",
          "content": "<p>I experience less boost to increase score with bigger models (my boost is about 0.025?), as your score increases, its likely that the smaller models will catch up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1717205,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2022-03-09T18:08:30.167000",
      "content": "<p>Could someone provide insights on the gain we can expect by climbing the ladder from effnet-b0 to effnet-b7 ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1716785,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-03-09T11:56:19.457000",
      "content": "<p>small model/image provide framework pipeline experiments, but definitely not good enough.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1716713": "Just like me I am sure, that other people here also are training their new approaches on a smaller test bench instead of a full fledged training to save massive time. How much score do  you gain after using bigger model or bigger image size?\n\nScores below will be updated as I improve\n\nLet me start with mine:\nB5 - 384 Single Fold (Usually the newest score, the later scores will be updated in a few days)\nCV: 0.696\nLB: 0.730\n\nB5 - 512 Single Fold\nCV: 0.707\nLB: 0.745\n\nB5 - 384 5 Folds\nLB: 0.767\n\nB5 - 512 5 folds\nLB: 0.778\n\nScores above will be updated as I improve",
    "1716793": "Yes, we can see correlation as well. \nIn our case we manage to increse score about 0.02-0.03 jumping from 384 -> 512. Certainly there is limitations we observe now:\n- GPU/TPU memory -> gradient exploding ... to small batch size ... to long training etc.\n- increasing size (for small object) there is chance to loose some important image features (or change them significantly) - we managed to find golden point in our dataset (still looking for improvements) which gave us 0.810. \n\nI thnink that many of us use the same approach here (TF ArcFace/KNN or Pytorch ArcFace/GeM). The difference is mainly in two parts:\n- dataset - we have different datasets (different ROI presented to NN). Moreover people from TOP3 (above 0.83) have better way to extract important features to NN.\n- inference - some trick / postprocessing - I am still thinking about using some trick with feature matching - notebook I published - [LoFTR notebook](https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia) (I tried SURF ans SIFT as well but they do not perform well in this situation).",
    "1717540": "model - > swin_base\nimage size -> 224\nLB = 0.512\n",
    "1718166": "batch_size=?",
    "1717903": "Your result is single fold or ensemble of kfold?",
    "1717493": "I've been using small models and img sizes to test out frameworks. \n\nB5 - 512\nDataset: Detic Crops\nLB: 0.695\n\nI suspect about a 0.04-0.07 increase from B5-B7 and larger image size.",
    "1717205": "Could someone provide insights on the gain we can expect by climbing the ladder from effnet-b0 to effnet-b7 ?\n\n ",
    "1716785": "small model/image provide framework pipeline experiments, but definitely not good enough."
  }
}