{
  "id": 396826,
  "title": "Dumb question about the fragments",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/396826",
  "author_name": "",
  "post_date": "2023-03-23T02:52:23.931374900Z",
  "votes": 3,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I'm beating my head against this, and I don't know if it's just been a long day at work without enough caffeine, but anyway …</p>\n<p>For the life of me, it looks like the two test fragments 'a' and 'b' are just the training fragment '1', cut horizontally 2,727 rows down from the top.</p>\n<p>Am I not comprehending the data files correctly? It looks the same no matter if I download the full zip file or just individual files.</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "2192974",
      "postDate": "03/23/2023 02:52:23",
      "content": "<p>I'm beating my head against this, and I don't know if it's just been a long day at work without enough caffeine, but anyway …</p>\n<p>For the life of me, it looks like the two test fragments 'a' and 'b' are just the training fragment '1', cut horizontally 2,727 rows down from the top.</p>\n<p>Am I not comprehending the data files correctly? It looks the same no matter if I download the full zip file or just individual files.</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "I'm beating my head against this, and I don't know if it's just been a long day at work without enough caffeine, but anyway …\n\nFor the life of me, it looks like the two test fragments 'a' and 'b' are just the training fragment '1', cut horizontally 2,727 rows down from the top.\n\nAm I not comprehending the data files correctly? It looks the same no matter if I download the full zip file or just individual files.\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "2193003",
      "postDate": "03/23/2023 03:25:53",
      "content": "<p>from Data <code>The sample slices available to download in the test folders are simply copied from training fragment one.</code>. Actual Test set is hidden.</p>",
      "rawMarkdown": "from Data ` The sample slices available to download in the test folders are simply copied from training fragment one.`. Actual Test set is hidden.",
      "votes": null
    },
    {
      "id": "2193006",
      "postDate": "03/23/2023 03:30:44",
      "content": "<p>The real 'a' and 'b' are hidden from you and your code will only see them when you do a submission.</p>\n<p>In many past competitions Kaggle provided similar 'fake' test data that was so limited in size and scope that you really could not evaluate the code inference very well.   It would seem they have realized that we need reasonable test data to avoid an endless string of discussion topics where folks are asking help because their code submission failed.</p>\n<p>Having a good size 'test' dataset will let you insure that your code can read more than one fragment, etc.</p>",
      "rawMarkdown": "The real 'a' and 'b' are hidden from you and your code will only see them when you do a submission.\n\nIn many past competitions Kaggle provided similar 'fake' test data that was so limited in size and scope that you really could not evaluate the code inference very well.   It would seem they have realized that we need reasonable test data to avoid an endless string of discussion topics where folks are asking help because their code submission failed.\n\nHaving a good size 'test' dataset will let you insure that your code can read more than one fragment, etc.",
      "votes": null
    },
    {
      "id": "2193007",
      "postDate": "03/23/2023 03:31:34",
      "content": "<p>Ah, thanks! I didn't read those closely enough. 😔</p>\n<p>So … if using a self-written program to do the training and testing, the Notebook will have to shell out to that program to run on the hidden test data? </p>\n<p>I need to understand these Notebooks better …&nbsp;I'm in the dark as to how I can compile a program on the box it's hosted on. Or maybe I'm not understanding that either.</p>",
      "rawMarkdown": "Ah, thanks! I didn't read those closely enough. 😔\n\nSo … if using a self-written program to do the training and testing, the Notebook will have to shell out to that program to run on the hidden test data? \n\nI need to understand these Notebooks better … I'm in the dark as to how I can compile a program on the box it's hosted on. Or maybe I'm not understanding that either.",
      "votes": null
    },
    {
      "id": "2193941",
      "postDate": "03/23/2023 15:53:34",
      "content": "<p>Yes, in the public dataset the test/ data is dummy data. When you submit your notebook it will be substituted for real data. This is to keep the test data secret. I've updated the data page to make this clearer, thanks for that.</p>",
      "rawMarkdown": "Yes, in the public dataset the test/ data is dummy data. When you submit your notebook it will be substituted for real data. This is to keep the test data secret. I've updated the data page to make this clearer, thanks for that.",
      "votes": null
    },
    {
      "id": "2194005",
      "postDate": "03/23/2023 16:56:39",
      "content": "<p>Thanks!</p>\n<p>So am I right that there is no way to compile code to run on your Jupyter host that could be shelled out to from a Notebook?</p>",
      "rawMarkdown": "Thanks!\n\nSo am I right that there is no way to compile code to run on your Jupyter host that could be shelled out to from a Notebook?",
      "votes": null
    },
    {
      "id": "2194032",
      "postDate": "03/23/2023 17:21:26",
      "content": "<p>I'm not sure exactly what can be run in a Jupyter notebook. I'll ask the Kaggle folks.</p>",
      "rawMarkdown": "I'm not sure exactly what can be run in a Jupyter notebook. I'll ask the Kaggle folks.",
      "votes": null
    },
    {
      "id": "2194085",
      "postDate": "03/23/2023 18:04:47",
      "content": "<p><a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a> if the code can be run on linux, you can upload it as a dataset and call it from within python, or using Juptyer's magic command <code>!</code></p>",
      "rawMarkdown": "costella if the code can be run on linux, you can upload it as a dataset and call it from within python, or using Juptyer's magic command `!`",
      "votes": null
    },
    {
      "id": "2194094",
      "postDate": "03/23/2023 18:09:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a>, if there are libraries or code that aren't in our Python or R environments, you could create a package and add it to <a href=\"https://www.kaggle.com/datasets\" target=\"_blank\">a dataset</a>, which you can then attach to your notebook. This would allow you to install the package without needing internet access turned on (not allowed during submission scoring). You could use a similar procedure for preprocessed data or pretrained models.</p>\n<p>Hope this answers your question!</p>",
      "rawMarkdown": "Hi @costella, if there are libraries or code that aren't in our Python or R environments, you could create a package and add it to [a dataset](https://www.kaggle.com/datasets), which you can then attach to your notebook. This would allow you to install the package without needing internet access turned on (not allowed during submission scoring). You could use a similar procedure for preprocessed data or pretrained models.\n\nHope this answers your question!",
      "votes": null
    },
    {
      "id": "2194105",
      "postDate": "03/23/2023 18:21:07",
      "content": "<p>Ah thanks all for these pointers! I'll look into this.</p>",
      "rawMarkdown": "Ah thanks all for these pointers! I'll look into this.",
      "votes": null
    },
    {
      "id": "2194108",
      "postDate": "03/23/2023 18:24:47",
      "content": "<p>So …&nbsp;could someone give me a pointer to where I can learn about creating a package to add to a dataset? I'm only familiar with Python packages, not compiled executables.</p>",
      "rawMarkdown": "So … could someone give me a pointer to where I can learn about creating a package to add to a dataset? I'm only familiar with Python packages, not compiled executables.",
      "votes": null
    },
    {
      "id": "2194110",
      "postDate": "03/23/2023 18:29:04",
      "content": "<p>If it's a regular executable, you should be able to (1) add the file(s) to a private dataset (2) attach the dataset to the notebook (3) make sure it's executable (4) call it from the notebook.</p>",
      "rawMarkdown": "If it's a regular executable, you should be able to (1) add the file(s) to a private dataset (2) attach the dataset to the notebook (3) make sure it's executable (4) call it from the notebook.",
      "votes": null
    },
    {
      "id": "2194118",
      "postDate": "03/23/2023 18:34:26",
      "content": "<p>But an executable generally needs to be compiled for the particular operating system it is being run on. I guess if you DM me with some details of that then I can compile it on a compatible OS and upload the executable directly?</p>",
      "rawMarkdown": "But an executable generally needs to be compiled for the particular operating system it is being run on. I guess if you DM me with some details of that then I can compile it on a compatible OS and upload the executable directly?",
      "votes": null
    },
    {
      "id": "2194134",
      "postDate": "03/23/2023 18:43:53",
      "content": "<p>You can also compile it on notebooks. It's rare that people do this (python is pretty much the name of the game these days), but there should be nothing stopping you from doing it. We've even seen folks do stuff like:</p>\n<pre><code>%%writefile my.cpp\n... C++ code ...\n\n!gcc ... my.cpp\n!./a.out \n</code></pre>",
      "rawMarkdown": "You can also compile it on notebooks. It's rare that people do this (python is pretty much the name of the game these days), but there should be nothing stopping you from doing it. We've even seen folks do stuff like:\n\n```\n%%writefile my.cpp\n... C++ code ...\n\n!gcc ... my.cpp\n!./a.out \n```",
      "votes": null
    },
    {
      "id": "2194147",
      "postDate": "03/23/2023 18:48:35",
      "content": "<p>It's not going to compile from within a notebook.</p>",
      "rawMarkdown": "It's not going to compile from within a notebook.",
      "votes": null
    },
    {
      "id": "2194856",
      "postDate": "03/24/2023 07:53:57",
      "content": "<p>OK, I didn't think about this long enough. I've figured it out now. 🙂 Thanks again!</p>",
      "rawMarkdown": "OK, I didn't think about this long enough. I've figured it out now. 🙂 Thanks again!",
      "votes": null
    },
    {
      "id": "2195062",
      "postDate": "03/24/2023 11:44:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a> <br>\nIt's possible that the test fragments 'a' and 'b' are indeed similar to the training fragment '1', but with some subtle differences. It's not uncommon for test data to be generated in a way that is similar to the training data, but with some variations to test the model's ability to generalize and make accurate predictions on new data.</p>",
      "rawMarkdown": "Hi @costella \nIt's possible that the test fragments 'a' and 'b' are indeed similar to the training fragment '1', but with some subtle differences. It's not uncommon for test data to be generated in a way that is similar to the training data, but with some variations to test the model's ability to generalize and make accurate predictions on new data.",
      "votes": null
    },
    {
      "id": "2195145",
      "postDate": "03/24/2023 12:44:34",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tariqbashir\" target=\"_blank\">@tariqbashir</a> . No, in this case I just didn't RTFM 😊 — the test data is just a mock, literally taken from the first training fragment by slicing it in two, as described above (and I think the Data section has had extra text added to it to make this even clearer). It's so that you can check that your Notebook is running properly with the mock test data. When you submit, I think they run it in a kernel with the real (secret) test fragments in place of the two mock fragments that are there now.</p>",
      "rawMarkdown": "Thanks @tariqbashir . No, in this case I just didn't RTFM 😊 — the test data is just a mock, literally taken from the first training fragment by slicing it in two, as described above (and I think the Data section has had extra text added to it to make this even clearer). It's so that you can check that your Notebook is running properly with the mock test data. When you submit, I think they run it in a kernel with the real (secret) test fragments in place of the two mock fragments that are there now.",
      "votes": null
    },
    {
      "id": "2213890",
      "postDate": "04/07/2023 23:47:23",
      "content": "<p>I am also confused about this. Is it possible to participate in this competition without signing up for a compute service? Is a compute service provided for free to the participants? If so where can I find instructions on using it? Can we effectively log in with a shell and author and test the code on the test samples?</p>\n<p>If not, is there some comprehensive list of available software and libraries? (I am used to arch linux)</p>\n<p>I'd hate to end up spending time working on code I might not be able to get to run by the time I wish to submit…</p>",
      "rawMarkdown": "I am also confused about this. Is it possible to participate in this competition without signing up for a compute service? Is a compute service provided for free to the participants? If so where can I find instructions on using it? Can we effectively log in with a shell and author and test the code on the test samples?\n\nIf not, is there some comprehensive list of available software and libraries? (I am used to arch linux)\n\nI'd hate to end up spending time working on code I might not be able to get to run by the time I wish to submit...",
      "votes": null
    },
    {
      "id": "2213998",
      "postDate": "04/08/2023 04:22:50",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ludwigmaes\" target=\"_blank\">@ludwigmaes</a>, I'm just figuring this out myself, but this is what I think I have figured out:</p>\n<ol>\n<li>You can download the training data. But none of the test set is directly downloadable.</li>\n<li>To make submissions you need to run in a Kaggle (Jupyter) Notebook. There are placeholders for the test data, which are just the first test fragment sliced in two, just so that you can get the execution flow right.</li>\n<li>When you submit, the Kaggle notebook is run on a subset of the actual test data, instead of the placeholders, and the results go onto the Leaderboard.</li>\n<li>You can't use the fancier Google cloud notebooks (whatever they're called) for submissions, so I'm avoiding them altogether.</li>\n<li>You can include arbitrary files in a private dataset that you can attach to your Kaggle Notebook.</li>\n<li>I have been able to compile C source code I uploaded in a dataset by shelling out single commands from the Jupyter notebook. Not simple, but doable.</li>\n</ol>",
      "rawMarkdown": "Hey @ludwigmaes, I'm just figuring this out myself, but this is what I think I have figured out:\n1. You can download the training data. But none of the test set is directly downloadable.\n2. To make submissions you need to run in a Kaggle (Jupyter) Notebook. There are placeholders for the test data, which are just the first test fragment sliced in two, just so that you can get the execution flow right.\n3. When you submit, the Kaggle notebook is run on a subset of the actual test data, instead of the placeholders, and the results go onto the Leaderboard.\n4. You can't use the fancier Google cloud notebooks (whatever they're called) for submissions, so I'm avoiding them altogether.\n5. You can include arbitrary files in a private dataset that you can attach to your Kaggle Notebook.\n6. I have been able to compile C source code I uploaded in a dataset by shelling out single commands from the Jupyter notebook. Not simple, but doable.",
      "votes": null
    },
    {
      "id": "2215044",
      "postDate": "04/09/2023 03:50:28",
      "content": "<p>Thanks for describing this, I'll link this in Discord for people who also want to run their own non-Python programs!</p>",
      "rawMarkdown": "Thanks for describing this, I'll link this in Discord for people who also want to run their own non-Python programs!",
      "votes": null
    },
    {
      "id": "2216549",
      "postDate": "04/10/2023 06:43:10",
      "content": "<p>Thank you, I never used the Kaggle Notebook environment before, it actually feels nice now that I tried it.</p>",
      "rawMarkdown": "Thank you, I never used the Kaggle Notebook environment before, it actually feels nice now that I tried it.",
      "votes": null
    },
    {
      "id": "2240578",
      "postDate": "04/30/2023 16:49:31",
      "content": "<p>Actually I am not sure if it's an issue with my code only but model that performs well on test/b data performs badly on test/a and both of which seems to be cropped versions of train/1 data. So maybe what <a href=\"https://www.kaggle.com/tariqbashir\" target=\"_blank\">@tariqbashir</a> said might be the case here.</p>",
      "rawMarkdown": "Actually I am not sure if it's an issue with my code only but model that performs well on test/b data performs badly on test/a and both of which seems to be cropped versions of train/1 data. So maybe what @tariqbashir said might be the case here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2193003,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "03/23/2023 03:25:53",
      "content": "<p>from Data <code>The sample slices available to download in the test folders are simply copied from training fragment one.</code>. Actual Test set is hidden.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2193007,
          "author_name": "costella",
          "author_url": "",
          "post_date": "03/23/2023 03:31:34",
          "content": "<p>Ah, thanks! I didn't read those closely enough. 😔</p>\n<p>So … if using a self-written program to do the training and testing, the Notebook will have to shell out to that program to run on the hidden test data? </p>\n<p>I need to understand these Notebooks better …&nbsp;I'm in the dark as to how I can compile a program on the box it's hosted on. Or maybe I'm not understanding that either.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2193006,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "03/23/2023 03:30:44",
      "content": "<p>The real 'a' and 'b' are hidden from you and your code will only see them when you do a submission.</p>\n<p>In many past competitions Kaggle provided similar 'fake' test data that was so limited in size and scope that you really could not evaluate the code inference very well.   It would seem they have realized that we need reasonable test data to avoid an endless string of discussion topics where folks are asking help because their code submission failed.</p>\n<p>Having a good size 'test' dataset will let you insure that your code can read more than one fragment, etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2193941,
      "author_name": "jpposma",
      "author_url": "",
      "post_date": "03/23/2023 15:53:34",
      "content": "<p>Yes, in the public dataset the test/ data is dummy data. When you submit your notebook it will be substituted for real data. This is to keep the test data secret. I've updated the data page to make this clearer, thanks for that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2194005,
          "author_name": "costella",
          "author_url": "",
          "post_date": "03/23/2023 16:56:39",
          "content": "<p>Thanks!</p>\n<p>So am I right that there is no way to compile code to run on your Jupyter host that could be shelled out to from a Notebook?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2194032,
              "author_name": "jpposma",
              "author_url": "",
              "post_date": "03/23/2023 17:21:26",
              "content": "<p>I'm not sure exactly what can be run in a Jupyter notebook. I'll ask the Kaggle folks.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2194085,
              "author_name": "wcukierski",
              "author_url": "",
              "post_date": "03/23/2023 18:04:47",
              "content": "<p><a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a> if the code can be run on linux, you can upload it as a dataset and call it from within python, or using Juptyer's magic command <code>!</code></p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2194094,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "03/23/2023 18:09:19",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a>, if there are libraries or code that aren't in our Python or R environments, you could create a package and add it to <a href=\"https://www.kaggle.com/datasets\" target=\"_blank\">a dataset</a>, which you can then attach to your notebook. This would allow you to install the package without needing internet access turned on (not allowed during submission scoring). You could use a similar procedure for preprocessed data or pretrained models.</p>\n<p>Hope this answers your question!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2194105,
                  "author_name": "costella",
                  "author_url": "",
                  "post_date": "03/23/2023 18:21:07",
                  "content": "<p>Ah thanks all for these pointers! I'll look into this.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2194108,
                      "author_name": "costella",
                      "author_url": "",
                      "post_date": "03/23/2023 18:24:47",
                      "content": "<p>So …&nbsp;could someone give me a pointer to where I can learn about creating a package to add to a dataset? I'm only familiar with Python packages, not compiled executables.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2194110,
                          "author_name": "wcukierski",
                          "author_url": "",
                          "post_date": "03/23/2023 18:29:04",
                          "content": "<p>If it's a regular executable, you should be able to (1) add the file(s) to a private dataset (2) attach the dataset to the notebook (3) make sure it's executable (4) call it from the notebook.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2194118,
                              "author_name": "costella",
                              "author_url": "",
                              "post_date": "03/23/2023 18:34:26",
                              "content": "<p>But an executable generally needs to be compiled for the particular operating system it is being run on. I guess if you DM me with some details of that then I can compile it on a compatible OS and upload the executable directly?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2194134,
                                  "author_name": "wcukierski",
                                  "author_url": "",
                                  "post_date": "03/23/2023 18:43:53",
                                  "content": "<p>You can also compile it on notebooks. It's rare that people do this (python is pretty much the name of the game these days), but there should be nothing stopping you from doing it. We've even seen folks do stuff like:</p>\n<pre><code>%%writefile my.cpp\n... C++ code ...\n\n!gcc ... my.cpp\n!./a.out \n</code></pre>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2194147,
                                      "author_name": "costella",
                                      "author_url": "",
                                      "post_date": "03/23/2023 18:48:35",
                                      "content": "<p>It's not going to compile from within a notebook.</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                },
                                {
                                  "id": 2194856,
                                  "author_name": "costella",
                                  "author_url": "",
                                  "post_date": "03/24/2023 07:53:57",
                                  "content": "<p>OK, I didn't think about this long enough. I've figured it out now. 🙂 Thanks again!</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2195062,
      "author_name": "tariqbashir",
      "author_url": "",
      "post_date": "03/24/2023 11:44:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/costella\" target=\"_blank\">@costella</a> <br>\nIt's possible that the test fragments 'a' and 'b' are indeed similar to the training fragment '1', but with some subtle differences. It's not uncommon for test data to be generated in a way that is similar to the training data, but with some variations to test the model's ability to generalize and make accurate predictions on new data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2195145,
          "author_name": "costella",
          "author_url": "",
          "post_date": "03/24/2023 12:44:34",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tariqbashir\" target=\"_blank\">@tariqbashir</a> . No, in this case I just didn't RTFM 😊 — the test data is just a mock, literally taken from the first training fragment by slicing it in two, as described above (and I think the Data section has had extra text added to it to make this even clearer). It's so that you can check that your Notebook is running properly with the mock test data. When you submit, I think they run it in a kernel with the real (secret) test fragments in place of the two mock fragments that are there now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2240578,
          "author_name": "vijaybj",
          "author_url": "",
          "post_date": "04/30/2023 16:49:31",
          "content": "<p>Actually I am not sure if it's an issue with my code only but model that performs well on test/b data performs badly on test/a and both of which seems to be cropped versions of train/1 data. So maybe what <a href=\"https://www.kaggle.com/tariqbashir\" target=\"_blank\">@tariqbashir</a> said might be the case here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2213890,
      "author_name": "ludwigmaes",
      "author_url": "",
      "post_date": "04/07/2023 23:47:23",
      "content": "<p>I am also confused about this. Is it possible to participate in this competition without signing up for a compute service? Is a compute service provided for free to the participants? If so where can I find instructions on using it? Can we effectively log in with a shell and author and test the code on the test samples?</p>\n<p>If not, is there some comprehensive list of available software and libraries? (I am used to arch linux)</p>\n<p>I'd hate to end up spending time working on code I might not be able to get to run by the time I wish to submit…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2213998,
          "author_name": "costella",
          "author_url": "",
          "post_date": "04/08/2023 04:22:50",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ludwigmaes\" target=\"_blank\">@ludwigmaes</a>, I'm just figuring this out myself, but this is what I think I have figured out:</p>\n<ol>\n<li>You can download the training data. But none of the test set is directly downloadable.</li>\n<li>To make submissions you need to run in a Kaggle (Jupyter) Notebook. There are placeholders for the test data, which are just the first test fragment sliced in two, just so that you can get the execution flow right.</li>\n<li>When you submit, the Kaggle notebook is run on a subset of the actual test data, instead of the placeholders, and the results go onto the Leaderboard.</li>\n<li>You can't use the fancier Google cloud notebooks (whatever they're called) for submissions, so I'm avoiding them altogether.</li>\n<li>You can include arbitrary files in a private dataset that you can attach to your Kaggle Notebook.</li>\n<li>I have been able to compile C source code I uploaded in a dataset by shelling out single commands from the Jupyter notebook. Not simple, but doable.</li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2215044,
              "author_name": "jpposma",
              "author_url": "",
              "post_date": "04/09/2023 03:50:28",
              "content": "<p>Thanks for describing this, I'll link this in Discord for people who also want to run their own non-Python programs!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2216549,
              "author_name": "ludwigmaes",
              "author_url": "",
              "post_date": "04/10/2023 06:43:10",
              "content": "<p>Thank you, I never used the Kaggle Notebook environment before, it actually feels nice now that I tried it.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2192974": "I'm beating my head against this, and I don't know if it's just been a long day at work without enough caffeine, but anyway …\n\nFor the life of me, it looks like the two test fragments 'a' and 'b' are just the training fragment '1', cut horizontally 2,727 rows down from the top.\n\nAm I not comprehending the data files correctly? It looks the same no matter if I download the full zip file or just individual files.\n\nThanks in advance!",
    "2193003": "from Data ` The sample slices available to download in the test folders are simply copied from training fragment one.`. Actual Test set is hidden.",
    "2193006": "The real 'a' and 'b' are hidden from you and your code will only see them when you do a submission.\n\nIn many past competitions Kaggle provided similar 'fake' test data that was so limited in size and scope that you really could not evaluate the code inference very well.   It would seem they have realized that we need reasonable test data to avoid an endless string of discussion topics where folks are asking help because their code submission failed.\n\nHaving a good size 'test' dataset will let you insure that your code can read more than one fragment, etc.",
    "2193007": "Ah, thanks! I didn't read those closely enough. 😔\n\nSo … if using a self-written program to do the training and testing, the Notebook will have to shell out to that program to run on the hidden test data? \n\nI need to understand these Notebooks better … I'm in the dark as to how I can compile a program on the box it's hosted on. Or maybe I'm not understanding that either.",
    "2193941": "Yes, in the public dataset the test/ data is dummy data. When you submit your notebook it will be substituted for real data. This is to keep the test data secret. I've updated the data page to make this clearer, thanks for that.",
    "2194005": "Thanks!\n\nSo am I right that there is no way to compile code to run on your Jupyter host that could be shelled out to from a Notebook?",
    "2194032": "I'm not sure exactly what can be run in a Jupyter notebook. I'll ask the Kaggle folks.",
    "2194085": "costella if the code can be run on linux, you can upload it as a dataset and call it from within python, or using Juptyer's magic command `!`",
    "2194094": "Hi @costella, if there are libraries or code that aren't in our Python or R environments, you could create a package and add it to [a dataset](https://www.kaggle.com/datasets), which you can then attach to your notebook. This would allow you to install the package without needing internet access turned on (not allowed during submission scoring). You could use a similar procedure for preprocessed data or pretrained models.\n\nHope this answers your question!",
    "2194105": "Ah thanks all for these pointers! I'll look into this.",
    "2194108": "So … could someone give me a pointer to where I can learn about creating a package to add to a dataset? I'm only familiar with Python packages, not compiled executables.",
    "2194110": "If it's a regular executable, you should be able to (1) add the file(s) to a private dataset (2) attach the dataset to the notebook (3) make sure it's executable (4) call it from the notebook.",
    "2194118": "But an executable generally needs to be compiled for the particular operating system it is being run on. I guess if you DM me with some details of that then I can compile it on a compatible OS and upload the executable directly?",
    "2194134": "You can also compile it on notebooks. It's rare that people do this (python is pretty much the name of the game these days), but there should be nothing stopping you from doing it. We've even seen folks do stuff like:\n\n```\n%%writefile my.cpp\n... C++ code ...\n\n!gcc ... my.cpp\n!./a.out \n```",
    "2194147": "It's not going to compile from within a notebook.",
    "2194856": "OK, I didn't think about this long enough. I've figured it out now. 🙂 Thanks again!",
    "2195062": "Hi @costella \nIt's possible that the test fragments 'a' and 'b' are indeed similar to the training fragment '1', but with some subtle differences. It's not uncommon for test data to be generated in a way that is similar to the training data, but with some variations to test the model's ability to generalize and make accurate predictions on new data.",
    "2195145": "Thanks @tariqbashir . No, in this case I just didn't RTFM 😊 — the test data is just a mock, literally taken from the first training fragment by slicing it in two, as described above (and I think the Data section has had extra text added to it to make this even clearer). It's so that you can check that your Notebook is running properly with the mock test data. When you submit, I think they run it in a kernel with the real (secret) test fragments in place of the two mock fragments that are there now.",
    "2213890": "I am also confused about this. Is it possible to participate in this competition without signing up for a compute service? Is a compute service provided for free to the participants? If so where can I find instructions on using it? Can we effectively log in with a shell and author and test the code on the test samples?\n\nIf not, is there some comprehensive list of available software and libraries? (I am used to arch linux)\n\nI'd hate to end up spending time working on code I might not be able to get to run by the time I wish to submit...",
    "2213998": "Hey @ludwigmaes, I'm just figuring this out myself, but this is what I think I have figured out:\n1. You can download the training data. But none of the test set is directly downloadable.\n2. To make submissions you need to run in a Kaggle (Jupyter) Notebook. There are placeholders for the test data, which are just the first test fragment sliced in two, just so that you can get the execution flow right.\n3. When you submit, the Kaggle notebook is run on a subset of the actual test data, instead of the placeholders, and the results go onto the Leaderboard.\n4. You can't use the fancier Google cloud notebooks (whatever they're called) for submissions, so I'm avoiding them altogether.\n5. You can include arbitrary files in a private dataset that you can attach to your Kaggle Notebook.\n6. I have been able to compile C source code I uploaded in a dataset by shelling out single commands from the Jupyter notebook. Not simple, but doable.",
    "2215044": "Thanks for describing this, I'll link this in Discord for people who also want to run their own non-Python programs!",
    "2216549": "Thank you, I never used the Kaggle Notebook environment before, it actually feels nice now that I tried it.",
    "2240578": "Actually I am not sure if it's an issue with my code only but model that performs well on test/b data performs badly on test/a and both of which seems to be cropped versions of train/1 data. So maybe what @tariqbashir said might be the case here."
  },
  "source": "meta"
}