{
  "id": 80003,
  "title": "Stage 1 vs. stage 2 kernel performance",
  "url": "/competitions/quora-insincere-questions-classification/discussion/80003",
  "author_name": "",
  "post_date": "2019-02-09T12:49:10.564077900Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Intro</h2>\n\n<p>Yesterday i was curious about my selected submissions so I ran both of them just to get some feeling about stage 2 performance and it seems like kernels are really slow! This could be due to kernel validation phase or is it something else? I tried to simulate stage 2 during first stage of the competition and I had plenty of time left.</p>\n\n<p>During stage 1 my final model had a <strong>mean(runtime)</strong> of ~ <strong>6400 seconds</strong>.</p>\n\n<h2>Test setup</h2>\n\n<p>During yesterdays (08/02/2019) stage 2 performance tests I commited two types of kernels. <strong>Kernel a)</strong> uses the full stage 2 test-set (375806 rows) and <strong>Kernel b)</strong> is based on the top 56370 rows (equal to stage 1 test size). </p>\n\n<p>I ran both kernels multiple times:</p>\n\n<ul>\n<li>Kernel a) stage 2 test-set (full = 375806 rows) with a mean(runtime) ~ <strong>7303</strong> seconds.</li>\n<li>Kernel b) stage 2 test-set (subset = 56370 rows) with a mean(runtime) ~ <strong>7002</strong> seconds.</li>\n</ul>\n\n<p>To rule out that this was not an isolated case I repeated <strong>Kernel b)</strong> today (09/02/2019):\nMin. = <strong>6861</strong> seconds\nMean = <strong>6939</strong> seconds \nMax. = <strong>7046</strong> seconds </p>\n\n<p>As you can see ... even the subsetted stage 2 version is ~ <strong>600</strong> seconds slower compared to stage 1's performance.</p>\n\n<p>I am using R by the way.</p>\n\n<h2>Kernel package versions changed?</h2>\n\n<p>After comparing stage 1 and 2 kernel dumps i found out that there has been some package changes inbetween stages\nStage 1:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11224/stage1.PNG\" alt=\"some stage 1 package info\">\nStage 2:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11225/stage2.PNG\" alt=\"some stage 2 package info\"></p>\n\n<p>Unfortunately I have no full/ complete SessionInfo of the stage 1 kernel...</p>\n\n<p>Have you noticed anything similar? </p>\n\n<p>Cheers\nDan</p>\n\n<h2>More testing</h2>\n\n<p>Did some testing over the weekend. Figured out that stage 2 preprocessing behaves different/ strange. That is why I created 6 test setups to compare preprocessing runtime:</p>\n\n<p><strong>Scenario 1 – sftp stage1 kernel:</strong>\nUploaded Stage 1 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Goal is to reproduce stage 1 environment. Removed file from sftp right after testing!</p>\n\n<p><strong>Scenario 2 – sftp stage 2 kernel:</strong>\nUploaded Stage 2 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Removed file from sftp right after testing!\n-   take a random sample of size 56370 </p>\n\n<p><strong>Scenario 3 – import stage 2 data hosted by kaggle</strong>\n-   take a random sample of size 56370 </p>\n\n<p><strong>Scenario 4 – stage 1 local computer:</strong>\nLike scenario 1 but on local computer.</p>\n\n<p><strong>Scenario 5 – stage 2 local computer</strong>\nLike scenario 2 but on local computer</p>\n\n<p><strong>Scenario 6 - full stage 2 data local computer</strong></p>\n\n<h2>Results</h2>\n\n<p><img src=\"https://storage.googleapis.com/dbi-share/elo/kaggle_stage2_timings.png\" alt=\"some test\"></p>\n\n<p>You can clearly see a huge gap between stage 1 and stage 2 kernel performance! Stage 2 <strong>R</strong> kernels are really slow and preprocessing time almost doubled! </p>\n\n<p>tldr:\nI tried to replicate stage 1 with subsetted stage 2 data.\nKernels seem to be slower compared to stage 1.\nKernel package versions changed.</p>",
  "messages": [
    {
      "id": "468680",
      "postDate": "02/09/2019 12:49:10",
      "content": "<h2>Intro</h2>\n\n<p>Yesterday i was curious about my selected submissions so I ran both of them just to get some feeling about stage 2 performance and it seems like kernels are really slow! This could be due to kernel validation phase or is it something else? I tried to simulate stage 2 during first stage of the competition and I had plenty of time left.</p>\n\n<p>During stage 1 my final model had a <strong>mean(runtime)</strong> of ~ <strong>6400 seconds</strong>.</p>\n\n<h2>Test setup</h2>\n\n<p>During yesterdays (08/02/2019) stage 2 performance tests I commited two types of kernels. <strong>Kernel a)</strong> uses the full stage 2 test-set (375806 rows) and <strong>Kernel b)</strong> is based on the top 56370 rows (equal to stage 1 test size). </p>\n\n<p>I ran both kernels multiple times:</p>\n\n<ul>\n<li>Kernel a) stage 2 test-set (full = 375806 rows) with a mean(runtime) ~ <strong>7303</strong> seconds.</li>\n<li>Kernel b) stage 2 test-set (subset = 56370 rows) with a mean(runtime) ~ <strong>7002</strong> seconds.</li>\n</ul>\n\n<p>To rule out that this was not an isolated case I repeated <strong>Kernel b)</strong> today (09/02/2019):\nMin. = <strong>6861</strong> seconds\nMean = <strong>6939</strong> seconds \nMax. = <strong>7046</strong> seconds </p>\n\n<p>As you can see ... even the subsetted stage 2 version is ~ <strong>600</strong> seconds slower compared to stage 1's performance.</p>\n\n<p>I am using R by the way.</p>\n\n<h2>Kernel package versions changed?</h2>\n\n<p>After comparing stage 1 and 2 kernel dumps i found out that there has been some package changes inbetween stages\nStage 1:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11224/stage1.PNG\" alt=\"some stage 1 package info\">\nStage 2:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11225/stage2.PNG\" alt=\"some stage 2 package info\"></p>\n\n<p>Unfortunately I have no full/ complete SessionInfo of the stage 1 kernel...</p>\n\n<p>Have you noticed anything similar? </p>\n\n<p>Cheers\nDan</p>\n\n<h2>More testing</h2>\n\n<p>Did some testing over the weekend. Figured out that stage 2 preprocessing behaves different/ strange. That is why I created 6 test setups to compare preprocessing runtime:</p>\n\n<p><strong>Scenario 1 – sftp stage1 kernel:</strong>\nUploaded Stage 1 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Goal is to reproduce stage 1 environment. Removed file from sftp right after testing!</p>\n\n<p><strong>Scenario 2 – sftp stage 2 kernel:</strong>\nUploaded Stage 2 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Removed file from sftp right after testing!\n-   take a random sample of size 56370 </p>\n\n<p><strong>Scenario 3 – import stage 2 data hosted by kaggle</strong>\n-   take a random sample of size 56370 </p>\n\n<p><strong>Scenario 4 – stage 1 local computer:</strong>\nLike scenario 1 but on local computer.</p>\n\n<p><strong>Scenario 5 – stage 2 local computer</strong>\nLike scenario 2 but on local computer</p>\n\n<p><strong>Scenario 6 - full stage 2 data local computer</strong></p>\n\n<h2>Results</h2>\n\n<p><img src=\"https://storage.googleapis.com/dbi-share/elo/kaggle_stage2_timings.png\" alt=\"some test\"></p>\n\n<p>You can clearly see a huge gap between stage 1 and stage 2 kernel performance! Stage 2 <strong>R</strong> kernels are really slow and preprocessing time almost doubled! </p>\n\n<p>tldr:\nI tried to replicate stage 1 with subsetted stage 2 data.\nKernels seem to be slower compared to stage 1.\nKernel package versions changed.</p>",
      "rawMarkdown": "## Intro ##\nYesterday i was curious about my selected submissions so I ran both of them just to get some feeling about stage 2 performance and it seems like kernels are really slow! This could be due to kernel validation phase or is it something else? I tried to simulate stage 2 during first stage of the competition and I had plenty of time left.\n\nDuring stage 1 my final model had a **mean(runtime)** of ~ **6400 seconds**.\n\n## Test setup ##\nDuring yesterdays (08/02/2019) stage 2 performance tests I commited two types of kernels. **Kernel a)** uses the full stage 2 test-set (375806 rows) and **Kernel b)** is based on the top 56370 rows (equal to stage 1 test size). \n\n\nI ran both kernels multiple times:\n\n- Kernel a) stage 2 test-set (full = 375806 rows) with a mean(runtime) ~ **7303** seconds.\n- Kernel b) stage 2 test-set (subset = 56370 rows) with a mean(runtime) ~ **7002** seconds.\n\nTo rule out that this was not an isolated case I repeated **Kernel b)** today (09/02/2019):\nMin. = **6861** seconds\nMean = **6939** seconds \nMax. = **7046** seconds \n\nAs you can see ... even the subsetted stage 2 version is ~ **600** seconds slower compared to stage 1's performance.\n\nI am using R by the way.\n\n## Kernel package versions changed?##\n\nAfter comparing stage 1 and 2 kernel dumps i found out that there has been some package changes inbetween stages\nStage 1:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11224/stage1.PNG\" alt=\"some stage 1 package info\">\nStage 2:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11225/stage2.PNG\" alt=\"some stage 2 package info\">\n\nUnfortunately I have no full/ complete SessionInfo of the stage 1 kernel...\n\nHave you noticed anything similar? \n\nCheers\nDan\n\n## More testing ##\nDid some testing over the weekend. Figured out that stage 2 preprocessing behaves different/ strange. That is why I created 6 test setups to compare preprocessing runtime:\n\n**Scenario 1 – sftp stage1 kernel:**\nUploaded Stage 1 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Goal is to reproduce stage 1 environment. Removed file from sftp right after testing!\n\n**Scenario 2 – sftp stage 2 kernel:**\nUploaded Stage 2 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Removed file from sftp right after testing!\n-\ttake a random sample of size 56370 \n\n**Scenario 3 – import stage 2 data hosted by kaggle**\n-\ttake a random sample of size 56370 \n\n**Scenario 4 – stage 1 local computer:**\nLike scenario 1 but on local computer.\n\n**Scenario 5 – stage 2 local computer**\nLike scenario 2 but on local computer\n\n**Scenario 6 - full stage 2 data local computer**\n\n## Results  ##\n<img src=\"https://storage.googleapis.com/dbi-share/elo/kaggle_stage2_timings.png\" alt=\"some test\">\n\nYou can clearly see a huge gap between stage 1 and stage 2 kernel performance! Stage 2 **R** kernels are really slow and preprocessing time almost doubled! \n\n\ntldr:\nI tried to replicate stage 1 with subsetted stage 2 data.\nKernels seem to be slower compared to stage 1.\nKernel package versions changed.",
      "votes": null
    },
    {
      "id": "468739",
      "postDate": "02/09/2019 15:15:52",
      "content": "<p>I replied to you already in the other thread, but I cannot see any runtime differences, but I am not using R. Maybe you are not padding sequences to a fixed length and longer sequences are present in new data making inference run slower? The first k rows might not necessarily be the public LB data. Or maybe something happening with pre-processing.</p>\n\n<p>Regardless of that, changing package versions / docker images inbetween stages should be a no-go.</p>",
      "rawMarkdown": "I replied to you already in the other thread, but I cannot see any runtime differences, but I am not using R. Maybe you are not padding sequences to a fixed length and longer sequences are present in new data making inference run slower? The first k rows might not necessarily be the public LB data. Or maybe something happening with pre-processing.\n\nRegardless of that, changing package versions / docker images inbetween stages should be a no-go.",
      "votes": null
    },
    {
      "id": "468852",
      "postDate": "02/09/2019 21:04:55",
      "content": "<p>Padding seq is fixed. I just figured out that one can change the docker-image version. Unfortunately all available versions &lt; today keep crashing... </p>",
      "rawMarkdown": "Padding seq is fixed. I just figured out that one can change the docker-image version. Unfortunately all available versions &lt; today keep crashing...",
      "votes": null
    },
    {
      "id": "469017",
      "postDate": "02/10/2019 10:10:55",
      "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a>, do you by any chance create a vocabulary using train+test data in your kernel?  Since the set of vocabularies is determined by Keras tokenizer, with the large stage-2 test data, Keras tokenizer may change the set of best vocabularies as well.</p>",
      "rawMarkdown": "springmanndaniel, do you by any chance create a vocabulary using train+test data in your kernel?  Since the set of vocabularies is determined by Keras tokenizer, with the large stage-2 test data, Keras tokenizer may change the set of best vocabularies as well.",
      "votes": null
    },
    {
      "id": "469095",
      "postDate": "02/10/2019 13:39:45",
      "content": "<p>Vocab is built with train data only and sequence length is fixed. </p>",
      "rawMarkdown": "Vocab is built with train data only and sequence length is fixed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 468739,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/09/2019 15:15:52",
      "content": "<p>I replied to you already in the other thread, but I cannot see any runtime differences, but I am not using R. Maybe you are not padding sequences to a fixed length and longer sequences are present in new data making inference run slower? The first k rows might not necessarily be the public LB data. Or maybe something happening with pre-processing.</p>\n\n<p>Regardless of that, changing package versions / docker images inbetween stages should be a no-go.</p>",
      "votes": null,
      "replies": [
        {
          "id": 468852,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/09/2019 21:04:55",
          "content": "<p>Padding seq is fixed. I just figured out that one can change the docker-image version. Unfortunately all available versions &lt; today keep crashing... </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469017,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "02/10/2019 10:10:55",
      "content": "<p><a href=\"/springmanndaniel\">@springmanndaniel</a>, do you by any chance create a vocabulary using train+test data in your kernel?  Since the set of vocabularies is determined by Keras tokenizer, with the large stage-2 test data, Keras tokenizer may change the set of best vocabularies as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 469095,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/10/2019 13:39:45",
          "content": "<p>Vocab is built with train data only and sequence length is fixed. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "468680": "## Intro ##\nYesterday i was curious about my selected submissions so I ran both of them just to get some feeling about stage 2 performance and it seems like kernels are really slow! This could be due to kernel validation phase or is it something else? I tried to simulate stage 2 during first stage of the competition and I had plenty of time left.\n\nDuring stage 1 my final model had a **mean(runtime)** of ~ **6400 seconds**.\n\n## Test setup ##\nDuring yesterdays (08/02/2019) stage 2 performance tests I commited two types of kernels. **Kernel a)** uses the full stage 2 test-set (375806 rows) and **Kernel b)** is based on the top 56370 rows (equal to stage 1 test size). \n\n\nI ran both kernels multiple times:\n\n- Kernel a) stage 2 test-set (full = 375806 rows) with a mean(runtime) ~ **7303** seconds.\n- Kernel b) stage 2 test-set (subset = 56370 rows) with a mean(runtime) ~ **7002** seconds.\n\nTo rule out that this was not an isolated case I repeated **Kernel b)** today (09/02/2019):\nMin. = **6861** seconds\nMean = **6939** seconds \nMax. = **7046** seconds \n\nAs you can see ... even the subsetted stage 2 version is ~ **600** seconds slower compared to stage 1's performance.\n\nI am using R by the way.\n\n## Kernel package versions changed?##\n\nAfter comparing stage 1 and 2 kernel dumps i found out that there has been some package changes inbetween stages\nStage 1:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11224/stage1.PNG\" alt=\"some stage 1 package info\">\nStage 2:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/468662/11225/stage2.PNG\" alt=\"some stage 2 package info\">\n\nUnfortunately I have no full/ complete SessionInfo of the stage 1 kernel...\n\nHave you noticed anything similar? \n\nCheers\nDan\n\n## More testing ##\nDid some testing over the weekend. Figured out that stage 2 preprocessing behaves different/ strange. That is why I created 6 test setups to compare preprocessing runtime:\n\n**Scenario 1 – sftp stage1 kernel:**\nUploaded Stage 1 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Goal is to reproduce stage 1 environment. Removed file from sftp right after testing!\n\n**Scenario 2 – sftp stage 2 kernel:**\nUploaded Stage 2 test-set to my private sftp (no public access), created a Kaggle-kernel that imports all datasets directly from this location. Removed file from sftp right after testing!\n-\ttake a random sample of size 56370 \n\n**Scenario 3 – import stage 2 data hosted by kaggle**\n-\ttake a random sample of size 56370 \n\n**Scenario 4 – stage 1 local computer:**\nLike scenario 1 but on local computer.\n\n**Scenario 5 – stage 2 local computer**\nLike scenario 2 but on local computer\n\n**Scenario 6 - full stage 2 data local computer**\n\n## Results  ##\n<img src=\"https://storage.googleapis.com/dbi-share/elo/kaggle_stage2_timings.png\" alt=\"some test\">\n\nYou can clearly see a huge gap between stage 1 and stage 2 kernel performance! Stage 2 **R** kernels are really slow and preprocessing time almost doubled! \n\n\ntldr:\nI tried to replicate stage 1 with subsetted stage 2 data.\nKernels seem to be slower compared to stage 1.\nKernel package versions changed.",
    "468739": "I replied to you already in the other thread, but I cannot see any runtime differences, but I am not using R. Maybe you are not padding sequences to a fixed length and longer sequences are present in new data making inference run slower? The first k rows might not necessarily be the public LB data. Or maybe something happening with pre-processing.\n\nRegardless of that, changing package versions / docker images inbetween stages should be a no-go.",
    "468852": "Padding seq is fixed. I just figured out that one can change the docker-image version. Unfortunately all available versions &lt; today keep crashing...",
    "469017": "springmanndaniel, do you by any chance create a vocabulary using train+test data in your kernel?  Since the set of vocabularies is determined by Keras tokenizer, with the large stage-2 test data, Keras tokenizer may change the set of best vocabularies as well.",
    "469095": "Vocab is built with train data only and sequence length is fixed."
  },
  "source": "meta"
}