{
  "id": 101327,
  "title": "How much are your two-site predictions consistent?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101327",
  "author_name": "",
  "post_date": "2019-07-25T01:52:51.763844200Z",
  "votes": 4,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I think it's a good strategy to run independent prediction for two sites and combine their result somehow, since it's the best form of data augmentation; data augmented in generation phase!</p>\n\n<p>Indeed, combining the two predictions gives me a significant improvement. However, I observed that the consistency of the two predictions are not as much as I expected.</p>\n\n<p>How about your case? Do you think training the net explicitly to have similar embeddings for the two images of a well will improve accuracy?</p>",
  "messages": [
    {
      "id": "583797",
      "postDate": "07/25/2019 01:52:51",
      "content": "<p>I think it's a good strategy to run independent prediction for two sites and combine their result somehow, since it's the best form of data augmentation; data augmented in generation phase!</p>\n\n<p>Indeed, combining the two predictions gives me a significant improvement. However, I observed that the consistency of the two predictions are not as much as I expected.</p>\n\n<p>How about your case? Do you think training the net explicitly to have similar embeddings for the two images of a well will improve accuracy?</p>",
      "rawMarkdown": "I think it's a good strategy to run independent prediction for two sites and combine their result somehow, since it's the best form of data augmentation; data augmented in generation phase!\n\nIndeed, combining the two predictions gives me a significant improvement. However, I observed that the consistency of the two predictions are not as much as I expected.\n\nHow about your case? Do you think training the net explicitly to have similar embeddings for the two images of a well will improve accuracy?",
      "votes": null
    },
    {
      "id": "584218",
      "postDate": "07/25/2019 15:17:04",
      "content": "<p>hi dohlee, thanks for sharing. are you training a specific network from scratch for each cell line? so using the particular classifier for the cell line, you train further on site-1 and site-2 and then predict both sites? thanks!</p>",
      "rawMarkdown": "hi dohlee, thanks for sharing. are you training a specific network from scratch for each cell line? so using the particular classifier for the cell line, you train further on site-1 and site-2 and then predict both sites? thanks!",
      "votes": null
    },
    {
      "id": "584299",
      "postDate": "07/25/2019 17:38:52",
      "content": "<p>Hi, so far, my model with the best LB score had imagenet-pretrained DenseNet121 as its backbone, which was first trained for the whole RxRx1 dataset. Using that imagenet+RxRx1-pretrained model, four models were then separately trained for each cell line. Of note, I used all the site-1 and site-2 images for training. Now I'm considering this as my baseline model, and I'm trying my ideas for the improvement. Hope this help!</p>",
      "rawMarkdown": "Hi, so far, my model with the best LB score had imagenet-pretrained DenseNet121 as its backbone, which was first trained for the whole RxRx1 dataset. Using that imagenet+RxRx1-pretrained model, four models were then separately trained for each cell line. Of note, I used all the site-1 and site-2 images for training. Now I'm considering this as my baseline model, and I'm trying my ideas for the improvement. Hope this help!",
      "votes": null
    },
    {
      "id": "584325",
      "postDate": "07/25/2019 18:30:15",
      "content": "<p>thanks dohlee! I am actually using the same approach but couldnt go further due to kernel limits. May i also know how many epochs are your training them? Thanks, once again! </p>",
      "rawMarkdown": "thanks dohlee! I am actually using the same approach but couldnt go further due to kernel limits. May i also know how many epochs are your training them? Thanks, once again!",
      "votes": null
    },
    {
      "id": "584329",
      "postDate": "07/25/2019 18:34:11",
      "content": "<p>With initial learning rate 1e-3, 25 epochs for pretraining model seemed best. Going further seemed to make my model overfit.</p>",
      "rawMarkdown": "With initial learning rate 1e-3, 25 epochs for pretraining model seemed best. Going further seemed to make my model overfit.",
      "votes": null
    },
    {
      "id": "584353",
      "postDate": "07/25/2019 19:47:18",
      "content": "<p><a href=\"/apap950419\">@apap950419</a>, I'm picking the site that has the highest confidence score across the two sites. For my best model (LB: 0.302), my predictions for site 1 and site 2 agree 45% of the time. </p>",
      "rawMarkdown": "apap950419, I'm picking the site that has the highest confidence score across the two sites. For my best model (LB: 0.302), my predictions for site 1 and site 2 agree 45% of the time.",
      "votes": null
    },
    {
      "id": "584356",
      "postDate": "07/25/2019 20:04:50",
      "content": "<p>Hi Anthony, in the event that I am not using AutoML - only <code>pytorch</code> and getting the  <code>CrossEntropyLoss</code>, what would \"confidence score\" translate to, for me?</p>",
      "rawMarkdown": "Hi Anthony, in the event that I am not using AutoML - only `pytorch` and getting the  `CrossEntropyLoss`, what would \"confidence score\" translate to, for me?",
      "votes": null
    },
    {
      "id": "584393",
      "postDate": "07/25/2019 21:39:03",
      "content": "<p>Hey <a href=\"/apap950419\">@apap950419</a> when do you say 'the whole RxRx1 dataset' do you mean the dataset in kaggle, or did you manage to find the original one?</p>\n\n<p>btw, I'm using the 2 sites like augmentation, randomly selecting one of them on each run, but I guess it shouldn't make much difference if you use both every time, maybe even better.</p>",
      "rawMarkdown": "Hey @apap950419 when do you say 'the whole RxRx1 dataset' do you mean the dataset in kaggle, or did you manage to find the original one?\n\nbtw, I'm using the 2 sites like augmentation, randomly selecting one of them on each run, but I guess it shouldn't make much difference if you use both every time, maybe even better.",
      "votes": null
    },
    {
      "id": "584416",
      "postDate": "07/25/2019 22:35:56",
      "content": "<p>i think he meant the current dataset on kaggle. that's already comprehensive. I didnt managet to use the whole dataset because im working on the kernels with a 9-hour time limit.</p>",
      "rawMarkdown": "i think he meant the current dataset on kaggle. that's already comprehensive. I didnt managet to use the whole dataset because im working on the kernels with a 9-hour time limit.",
      "votes": null
    },
    {
      "id": "584418",
      "postDate": "07/25/2019 22:41:33",
      "content": "<p>hey dohlee, are you also using adam/sgd optimizer or any lr schduler in specific to help convergence? wonder why mine is so slow with adam <code>lr=3e-4</code> and <code>exponential lr_scheduler</code>. </p>\n\n<p>also, are you training the entire training set during the first phase? Or how are you doing your validation for training on the whole dataset?</p>\n\n<p>thanks once again!</p>",
      "rawMarkdown": "hey dohlee, are you also using adam/sgd optimizer or any lr schduler in specific to help convergence? wonder why mine is so slow with adam `lr=3e-4` and `exponential lr_scheduler`. \n\nalso, are you training the entire training set during the first phase? Or how are you doing your validation for training on the whole dataset?\n\nthanks once again!",
      "votes": null
    },
    {
      "id": "584567",
      "postDate": "07/26/2019 07:09:33",
      "content": "<p>By slow you mean the loss decreases slowly?</p>",
      "rawMarkdown": "By slow you mean the loss decreases slowly?",
      "votes": null
    },
    {
      "id": "584794",
      "postDate": "07/26/2019 13:47:46",
      "content": "<p>hi lorenzo, yes, it decreases very slowly.</p>",
      "rawMarkdown": "hi lorenzo, yes, it decreases very slowly.",
      "votes": null
    },
    {
      "id": "584863",
      "postDate": "07/26/2019 15:38:14",
      "content": "<p>Not sure. Sorry. </p>",
      "rawMarkdown": "Not sure. Sorry.",
      "votes": null
    },
    {
      "id": "585205",
      "postDate": "07/27/2019 05:36:17",
      "content": "<p>How about trying some larger learning rate? I'm using Adam and multistepLR scheduler.</p>",
      "rawMarkdown": "How about trying some larger learning rate? I'm using Adam and multistepLR scheduler.",
      "votes": null
    },
    {
      "id": "585233",
      "postDate": "07/27/2019 06:24:03",
      "content": "<p>Hi <a href=\"/wjshenggggg\">@wjshenggggg</a> , it means your direct softmax/sigmoid output (each output shows probability [softmax case] or confidence [sigmoid case] of each class)</p>",
      "rawMarkdown": "Hi @wjshenggggg , it means your direct softmax/sigmoid output (each output shows probability [softmax case] or confidence [sigmoid case] of each class)",
      "votes": null
    },
    {
      "id": "585476",
      "postDate": "07/27/2019 14:22:00",
      "content": "<p>hi <a href=\"/ratthachat\">@ratthachat</a>, please correct me if i am wrong. softmax is rather the \"probability\" across all classes. say if it predicts the label 0, means label 0 has the highest probability among all 1108 classes. In our case, if label 0 of an <code>id_code</code> in site 1 has the highest softmax but label 2 of the same <code>id_code</code> has the highest softmax, we cannot directly consider the case where both sites \"agree 45% of the time\", can we? </p>",
      "rawMarkdown": "hi @ratthachat, please correct me if i am wrong. softmax is rather the \"probability\" across all classes. say if it predicts the label 0, means label 0 has the highest probability among all 1108 classes. In our case, if label 0 of an `id_code` in site 1 has the highest softmax but label 2 of the same `id_code` has the highest softmax, we cannot directly consider the case where both sites \"agree 45% of the time\", can we?",
      "votes": null
    },
    {
      "id": "585477",
      "postDate": "07/27/2019 14:23:27",
      "content": "<p>thanks <a href=\"/apap950419\">@apap950419</a>! will try them out! may i also ask about your validation strategy? I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case of <code>HEPG2</code>, it's\n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>",
      "rawMarkdown": "thanks @apap950419! will try them out! may i also ask about your validation strategy? I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case of `HEPG2`, it's\n```\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```",
      "votes": null
    },
    {
      "id": "585631",
      "postDate": "07/27/2019 19:05:10",
      "content": "<p>I'm simply using HEPG2-01, HUVEC-01, RPE-01, U2OS-01 as my validation set.</p>",
      "rawMarkdown": "I'm simply using HEPG2-01, HUVEC-01, RPE-01, U2OS-01 as my validation set.",
      "votes": null
    },
    {
      "id": "585712",
      "postDate": "07/27/2019 22:29:25",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> , yes, I think your statement is correct. Hopefully, I understand correctly. \nI think Anthony just meant , 45% of the time, his model predict the same highest-probability label for site1 and site2.</p>",
      "rawMarkdown": "wjshenggggg , yes, I think your statement is correct. Hopefully, I understand correctly. \nI think Anthony just meant , 45% of the time, his model predict the same highest-probability label for site1 and site2.",
      "votes": null
    },
    {
      "id": "585713",
      "postDate": "07/27/2019 22:33:22",
      "content": "<p>oh thank you! i thought he was able to select a site with higher confidence when both sites do not agree. so in that case, i wonder if comparison between both sites of the same <code>id_code</code> would actually make sense.</p>",
      "rawMarkdown": "oh thank you! i thought he was able to select a site with higher confidence when both sites do not agree. so in that case, i wonder if comparison between both sites of the same `id_code` would actually make sense.",
      "votes": null
    },
    {
      "id": "585952",
      "postDate": "07/28/2019 11:40:18",
      "content": "<p><a href=\"/apap950419\">@apap950419</a>  thanks for the tips!\nhow does your lb score differs from the validation score? \ncheers</p>",
      "rawMarkdown": "apap950419  thanks for the tips!\nhow does your lb score differs from the validation score? \ncheers",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 584218,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "07/25/2019 15:17:04",
      "content": "<p>hi dohlee, thanks for sharing. are you training a specific network from scratch for each cell line? so using the particular classifier for the cell line, you train further on site-1 and site-2 and then predict both sites? thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 584299,
          "author_name": "apap950419",
          "author_url": "",
          "post_date": "07/25/2019 17:38:52",
          "content": "<p>Hi, so far, my model with the best LB score had imagenet-pretrained DenseNet121 as its backbone, which was first trained for the whole RxRx1 dataset. Using that imagenet+RxRx1-pretrained model, four models were then separately trained for each cell line. Of note, I used all the site-1 and site-2 images for training. Now I'm considering this as my baseline model, and I'm trying my ideas for the improvement. Hope this help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584325,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/25/2019 18:30:15",
          "content": "<p>thanks dohlee! I am actually using the same approach but couldnt go further due to kernel limits. May i also know how many epochs are your training them? Thanks, once again! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584329,
          "author_name": "apap950419",
          "author_url": "",
          "post_date": "07/25/2019 18:34:11",
          "content": "<p>With initial learning rate 1e-3, 25 epochs for pretraining model seemed best. Going further seemed to make my model overfit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584393,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/25/2019 21:39:03",
          "content": "<p>Hey <a href=\"/apap950419\">@apap950419</a> when do you say 'the whole RxRx1 dataset' do you mean the dataset in kaggle, or did you manage to find the original one?</p>\n\n<p>btw, I'm using the 2 sites like augmentation, randomly selecting one of them on each run, but I guess it shouldn't make much difference if you use both every time, maybe even better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584416,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/25/2019 22:35:56",
          "content": "<p>i think he meant the current dataset on kaggle. that's already comprehensive. I didnt managet to use the whole dataset because im working on the kernels with a 9-hour time limit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584418,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/25/2019 22:41:33",
          "content": "<p>hey dohlee, are you also using adam/sgd optimizer or any lr schduler in specific to help convergence? wonder why mine is so slow with adam <code>lr=3e-4</code> and <code>exponential lr_scheduler</code>. </p>\n\n<p>also, are you training the entire training set during the first phase? Or how are you doing your validation for training on the whole dataset?</p>\n\n<p>thanks once again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584567,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "07/26/2019 07:09:33",
          "content": "<p>By slow you mean the loss decreases slowly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584794,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/26/2019 13:47:46",
          "content": "<p>hi lorenzo, yes, it decreases very slowly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585205,
          "author_name": "apap950419",
          "author_url": "",
          "post_date": "07/27/2019 05:36:17",
          "content": "<p>How about trying some larger learning rate? I'm using Adam and multistepLR scheduler.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585477,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 14:23:27",
          "content": "<p>thanks <a href=\"/apap950419\">@apap950419</a>! will try them out! may i also ask about your validation strategy? I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case of <code>HEPG2</code>, it's\n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585631,
          "author_name": "apap950419",
          "author_url": "",
          "post_date": "07/27/2019 19:05:10",
          "content": "<p>I'm simply using HEPG2-01, HUVEC-01, RPE-01, U2OS-01 as my validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585952,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/28/2019 11:40:18",
          "content": "<p><a href=\"/apap950419\">@apap950419</a>  thanks for the tips!\nhow does your lb score differs from the validation score? \ncheers</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 584353,
      "author_name": "antgoldbloom",
      "author_url": "",
      "post_date": "07/25/2019 19:47:18",
      "content": "<p><a href=\"/apap950419\">@apap950419</a>, I'm picking the site that has the highest confidence score across the two sites. For my best model (LB: 0.302), my predictions for site 1 and site 2 agree 45% of the time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 584356,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/25/2019 20:04:50",
          "content": "<p>Hi Anthony, in the event that I am not using AutoML - only <code>pytorch</code> and getting the  <code>CrossEntropyLoss</code>, what would \"confidence score\" translate to, for me?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584863,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/26/2019 15:38:14",
          "content": "<p>Not sure. Sorry. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585233,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/27/2019 06:24:03",
          "content": "<p>Hi <a href=\"/wjshenggggg\">@wjshenggggg</a> , it means your direct softmax/sigmoid output (each output shows probability [softmax case] or confidence [sigmoid case] of each class)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585476,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 14:22:00",
          "content": "<p>hi <a href=\"/ratthachat\">@ratthachat</a>, please correct me if i am wrong. softmax is rather the \"probability\" across all classes. say if it predicts the label 0, means label 0 has the highest probability among all 1108 classes. In our case, if label 0 of an <code>id_code</code> in site 1 has the highest softmax but label 2 of the same <code>id_code</code> has the highest softmax, we cannot directly consider the case where both sites \"agree 45% of the time\", can we? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585712,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/27/2019 22:29:25",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> , yes, I think your statement is correct. Hopefully, I understand correctly. \nI think Anthony just meant , 45% of the time, his model predict the same highest-probability label for site1 and site2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585713,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 22:33:22",
          "content": "<p>oh thank you! i thought he was able to select a site with higher confidence when both sites do not agree. so in that case, i wonder if comparison between both sites of the same <code>id_code</code> would actually make sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "583797": "I think it's a good strategy to run independent prediction for two sites and combine their result somehow, since it's the best form of data augmentation; data augmented in generation phase!\n\nIndeed, combining the two predictions gives me a significant improvement. However, I observed that the consistency of the two predictions are not as much as I expected.\n\nHow about your case? Do you think training the net explicitly to have similar embeddings for the two images of a well will improve accuracy?",
    "584218": "hi dohlee, thanks for sharing. are you training a specific network from scratch for each cell line? so using the particular classifier for the cell line, you train further on site-1 and site-2 and then predict both sites? thanks!",
    "584299": "Hi, so far, my model with the best LB score had imagenet-pretrained DenseNet121 as its backbone, which was first trained for the whole RxRx1 dataset. Using that imagenet+RxRx1-pretrained model, four models were then separately trained for each cell line. Of note, I used all the site-1 and site-2 images for training. Now I'm considering this as my baseline model, and I'm trying my ideas for the improvement. Hope this help!",
    "584325": "thanks dohlee! I am actually using the same approach but couldnt go further due to kernel limits. May i also know how many epochs are your training them? Thanks, once again!",
    "584329": "With initial learning rate 1e-3, 25 epochs for pretraining model seemed best. Going further seemed to make my model overfit.",
    "584353": "apap950419, I'm picking the site that has the highest confidence score across the two sites. For my best model (LB: 0.302), my predictions for site 1 and site 2 agree 45% of the time.",
    "584356": "Hi Anthony, in the event that I am not using AutoML - only `pytorch` and getting the  `CrossEntropyLoss`, what would \"confidence score\" translate to, for me?",
    "584393": "Hey @apap950419 when do you say 'the whole RxRx1 dataset' do you mean the dataset in kaggle, or did you manage to find the original one?\n\nbtw, I'm using the 2 sites like augmentation, randomly selecting one of them on each run, but I guess it shouldn't make much difference if you use both every time, maybe even better.",
    "584416": "i think he meant the current dataset on kaggle. that's already comprehensive. I didnt managet to use the whole dataset because im working on the kernels with a 9-hour time limit.",
    "584418": "hey dohlee, are you also using adam/sgd optimizer or any lr schduler in specific to help convergence? wonder why mine is so slow with adam `lr=3e-4` and `exponential lr_scheduler`. \n\nalso, are you training the entire training set during the first phase? Or how are you doing your validation for training on the whole dataset?\n\nthanks once again!",
    "584567": "By slow you mean the loss decreases slowly?",
    "584794": "hi lorenzo, yes, it decreases very slowly.",
    "584863": "Not sure. Sorry.",
    "585205": "How about trying some larger learning rate? I'm using Adam and multistepLR scheduler.",
    "585233": "Hi @wjshenggggg , it means your direct softmax/sigmoid output (each output shows probability [softmax case] or confidence [sigmoid case] of each class)",
    "585476": "hi @ratthachat, please correct me if i am wrong. softmax is rather the \"probability\" across all classes. say if it predicts the label 0, means label 0 has the highest probability among all 1108 classes. In our case, if label 0 of an `id_code` in site 1 has the highest softmax but label 2 of the same `id_code` has the highest softmax, we cannot directly consider the case where both sites \"agree 45% of the time\", can we?",
    "585477": "thanks @apap950419! will try them out! may i also ask about your validation strategy? I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case of `HEPG2`, it's\n```\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```",
    "585631": "I'm simply using HEPG2-01, HUVEC-01, RPE-01, U2OS-01 as my validation set.",
    "585712": "wjshenggggg , yes, I think your statement is correct. Hopefully, I understand correctly. \nI think Anthony just meant , 45% of the time, his model predict the same highest-probability label for site1 and site2.",
    "585713": "oh thank you! i thought he was able to select a site with higher confidence when both sites do not agree. so in that case, i wonder if comparison between both sites of the same `id_code` would actually make sense.",
    "585952": "apap950419  thanks for the tips!\nhow does your lb score differs from the validation score? \ncheers"
  },
  "source": "meta"
}