{
  "id": 217766,
  "title": "Do we need to get reproducible results?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217766",
  "author_name": "Rohan Suri",
  "post_date": "2021-02-08T10:34:38.442000",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This is a really basic question but do we need to get exact reproducible results in image competitions especially when using neural networks. If yes, how?<br>\nThanks.</p>",
  "messages": [
    {
      "id": 1191227,
      "postDate": "2021-02-08T10:34:38.443Z",
      "content": "<p>This is a really basic question but do we need to get exact reproducible results in image competitions especially when using neural networks. If yes, how?<br>\nThanks.</p>",
      "rawMarkdown": "This is a really basic question but do we need to get exact reproducible results in image competitions especially when using neural networks. If yes, how?\nThanks.",
      "votes": 3
    },
    {
      "id": 1191840,
      "postDate": "2021-02-08T17:58:17.837Z",
      "content": "<p>Yes.  You do not have to ensure it but keep the reproducibility of each experiments always helps. Some are listed below.</p>\n<ol>\n<li>Treat seed as a variable and evaluate the influence of random factor while comparing runs.</li>\n<li>Debug and test the correctness of the training pipeline with reproducible output.</li>\n</ol>",
      "rawMarkdown": "Yes.  You do not have to ensure it but keep the reproducibility of each experiments always helps. Some are listed below.\n1. Treat seed as a variable and evaluate the influence of random factor while comparing runs.\n2. Debug and test the correctness of the training pipeline with reproducible output.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1191851,
          "postDate": "2021-02-08T18:10:06.053Z",
          "content": "<p>I've tried setting up a seed value, but I still get slighly variable results in keras everytime I retrain. </p>",
          "rawMarkdown": "I've tried setting up a seed value, but I still get slighly variable results in keras everytime I retrain. "
        },
        {
          "id": 1193139,
          "postDate": "2021-02-09T13:35:55.200Z",
          "content": "<p>I'd firmly recommend you to chill. 😁</p>",
          "rawMarkdown": "I'd firmly recommend you to chill. 😁"
        },
        {
          "id": 1194704,
          "postDate": "2021-02-10T10:06:04.317Z",
          "content": "<p>hahaha sure✌️ </p>",
          "rawMarkdown": "hahaha sure✌️ ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1191615,
      "postDate": "2021-02-08T15:06:08.550Z",
      "content": "<p>Technically, it is not mandatory for you to have reproducible results. You just need an inference notebook which you can submit to a competition. </p>",
      "rawMarkdown": "Technically, it is not mandatory for you to have reproducible results. You just need an inference notebook which you can submit to a competition. ",
      "replies": [
        {
          "id": 1191726,
          "postDate": "2021-02-08T16:23:04.267Z",
          "content": "<p>When I say reproducible results, I mean that in one program run you get an accuracy of (let's say) 79.91 and in the other program run you get an accuracy of 79.92. Second accuracy is slightly different from the first run. Especially when using neural net, accuracy varies slightly everytime you run it from scratch. So, my question is if such variations are accepted or if there is a way to handle it , and if that handling is at all required?.<br>\nThanks.</p>",
          "rawMarkdown": "When I say reproducible results, I mean that in one program run you get an accuracy of (let's say) 79.91 and in the other program run you get an accuracy of 79.92. Second accuracy is slightly different from the first run. Especially when using neural net, accuracy varies slightly everytime you run it from scratch. So, my question is if such variations are accepted or if there is a way to handle it , and if that handling is at all required?.\nThanks."
        },
        {
          "id": 1191746,
          "postDate": "2021-02-08T16:41:15.127Z",
          "content": "<p>I don't think that's necessary, it's up to the user. As long as the scores are similar, both programs should be valid.</p>",
          "rawMarkdown": "I don't think that's necessary, it's up to the user. As long as the scores are similar, both programs should be valid.",
          "votes": 1
        },
        {
          "id": 1191831,
          "postDate": "2021-02-08T17:52:39.017Z",
          "content": "<p>I am still not clear about what exactly you are asking. All that matters is the final submission score that Kaggle shows. The score you get when you run it locally does not matter.</p>",
          "rawMarkdown": "I am still not clear about what exactly you are asking. All that matters is the final submission score that Kaggle shows. The score you get when you run it locally does not matter."
        },
        {
          "id": 1191838,
          "postDate": "2021-02-08T17:57:52.047Z",
          "content": "<p>Suppose you create a CNN model using Keras. Then everytime you rerun the same notebook, you might get slightly different results due to the randomization in weights in the training process. I was asking if that is an issue or not.</p>",
          "rawMarkdown": "Suppose you create a CNN model using Keras. Then everytime you rerun the same notebook, you might get slightly different results due to the randomization in weights in the training process. I was asking if that is an issue or not."
        },
        {
          "id": 1191844,
          "postDate": "2021-02-08T18:00:31.753Z",
          "content": "<p>No it's not, you can submit any version you like. And most of the time you'll be close to the maximum possible accuracy for a given model.</p>",
          "rawMarkdown": "No it's not, you can submit any version you like. And most of the time you'll be close to the maximum possible accuracy for a given model.",
          "votes": 1
        },
        {
          "id": 1193071,
          "postDate": "2021-02-09T12:46:54.007Z",
          "content": "<p>Any differences like you describe are going to be insignificant compared to the difference in score you get when you actually submit to Kaggle, so small variations in your local score should really not something to worry about.</p>",
          "rawMarkdown": "Any differences like you describe are going to be insignificant compared to the difference in score you get when you actually submit to Kaggle, so small variations in your local score should really not something to worry about.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1191604,
      "postDate": "2021-02-08T15:02:55.120Z",
      "content": "<p>But, in the Kaggle community, it is considered a good practice if your kernel has public or accesible data sources and has code written from scratch. It displays your honest work and improves transparency. Such notebooks are highly appreciated, even if their accuracy is relatively low. Datasets should follow the same pattern as far as possible. </p>",
      "rawMarkdown": "But, in the Kaggle community, it is considered a good practice if your kernel has public or accesible data sources and has code written from scratch. It displays your honest work and improves transparency. Such notebooks are highly appreciated, even if their accuracy is relatively low. Datasets should follow the same pattern as far as possible. "
    },
    {
      "id": 1191713,
      "postDate": "2021-02-08T16:14:29.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1191840,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2021-02-08T17:58:17.837000",
      "content": "<p>Yes.  You do not have to ensure it but keep the reproducibility of each experiments always helps. Some are listed below.</p>\n<ol>\n<li>Treat seed as a variable and evaluate the influence of random factor while comparing runs.</li>\n<li>Debug and test the correctness of the training pipeline with reproducible output.</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1191851,
          "author_name": "Rohan Suri",
          "author_url": "",
          "post_date": "2021-02-08T18:10:06.053000",
          "content": "<p>I've tried setting up a seed value, but I still get slighly variable results in keras everytime I retrain. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1193139,
          "author_name": "Aditya Kane",
          "author_url": "",
          "post_date": "2021-02-09T13:35:55.200000",
          "content": "<p>I'd firmly recommend you to chill. 😁</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1194704,
          "author_name": "Rohan Suri",
          "author_url": "",
          "post_date": "2021-02-10T10:06:04.317000",
          "content": "<p>hahaha sure✌️ </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1191615,
      "author_name": "Aditya Kane",
      "author_url": "",
      "post_date": "2021-02-08T15:06:08.550000",
      "content": "<p>Technically, it is not mandatory for you to have reproducible results. You just need an inference notebook which you can submit to a competition. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1191726,
          "author_name": "Rohan Suri",
          "author_url": "",
          "post_date": "2021-02-08T16:23:04.267000",
          "content": "<p>When I say reproducible results, I mean that in one program run you get an accuracy of (let's say) 79.91 and in the other program run you get an accuracy of 79.92. Second accuracy is slightly different from the first run. Especially when using neural net, accuracy varies slightly everytime you run it from scratch. So, my question is if such variations are accepted or if there is a way to handle it , and if that handling is at all required?.<br>\nThanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1191746,
          "author_name": "Aditya Kane",
          "author_url": "",
          "post_date": "2021-02-08T16:41:15.127000",
          "content": "<p>I don't think that's necessary, it's up to the user. As long as the scores are similar, both programs should be valid.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1191831,
          "author_name": "impulsecorp",
          "author_url": "",
          "post_date": "2021-02-08T17:52:39.017000",
          "content": "<p>I am still not clear about what exactly you are asking. All that matters is the final submission score that Kaggle shows. The score you get when you run it locally does not matter.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1191838,
          "author_name": "Rohan Suri",
          "author_url": "",
          "post_date": "2021-02-08T17:57:52.047000",
          "content": "<p>Suppose you create a CNN model using Keras. Then everytime you rerun the same notebook, you might get slightly different results due to the randomization in weights in the training process. I was asking if that is an issue or not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1191844,
          "author_name": "Aditya Kane",
          "author_url": "",
          "post_date": "2021-02-08T18:00:31.753000",
          "content": "<p>No it's not, you can submit any version you like. And most of the time you'll be close to the maximum possible accuracy for a given model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1193071,
          "author_name": "impulsecorp",
          "author_url": "",
          "post_date": "2021-02-09T12:46:54.007000",
          "content": "<p>Any differences like you describe are going to be insignificant compared to the difference in score you get when you actually submit to Kaggle, so small variations in your local score should really not something to worry about.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1191604,
      "author_name": "Aditya Kane",
      "author_url": "",
      "post_date": "2021-02-08T15:02:55.120000",
      "content": "<p>But, in the Kaggle community, it is considered a good practice if your kernel has public or accesible data sources and has code written from scratch. It displays your honest work and improves transparency. Such notebooks are highly appreciated, even if their accuracy is relatively low. Datasets should follow the same pattern as far as possible. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1191713,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-08T16:14:29.833000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1191227": "This is a really basic question but do we need to get exact reproducible results in image competitions especially when using neural networks. If yes, how?\nThanks.",
    "1191840": "Yes.  You do not have to ensure it but keep the reproducibility of each experiments always helps. Some are listed below.\n1. Treat seed as a variable and evaluate the influence of random factor while comparing runs.\n2. Debug and test the correctness of the training pipeline with reproducible output.\n\n",
    "1191615": "Technically, it is not mandatory for you to have reproducible results. You just need an inference notebook which you can submit to a competition. ",
    "1191604": "But, in the Kaggle community, it is considered a good practice if your kernel has public or accesible data sources and has code written from scratch. It displays your honest work and improves transparency. Such notebooks are highly appreciated, even if their accuracy is relatively low. Datasets should follow the same pattern as far as possible. ",
    "1191713": ""
  }
}