{
  "id": 490619,
  "title": "Get started here!",
  "url": "/competitions/leash-BELKA/discussion/490619",
  "author_name": "",
  "post_date": "2024-04-03T01:18:46.462107900Z",
  "votes": 15,
  "comment_count": 24,
  "views": 0,
  "content": "<p><strong>New to machine learning and data science?</strong> No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p><strong>New to Kaggle?</strong> Take a look at a few videos to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>. Publish and share your <a href=\"https://www.kaggle.com/docs/models#publishing-a-model\" target=\"_blank\">models on Kaggle Models</a>!</p>\n<p><strong>Looking for a team?</strong> Express your interest in joining a team through our <a href=\"https://www.kaggle.com/discussions/product-feedback/341195\" target=\"_blank\">Team Up</a> feature .</p>\n<p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our Kaggle community guidelines.</p>",
  "messages": [
    {
      "id": "2732053",
      "postDate": "04/03/2024 01:18:46",
      "content": "<p><strong>New to machine learning and data science?</strong> No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p><strong>New to Kaggle?</strong> Take a look at a few videos to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>. Publish and share your <a href=\"https://www.kaggle.com/docs/models#publishing-a-model\" target=\"_blank\">models on Kaggle Models</a>!</p>\n<p><strong>Looking for a team?</strong> Express your interest in joining a team through our <a href=\"https://www.kaggle.com/discussions/product-feedback/341195\" target=\"_blank\">Team Up</a> feature .</p>\n<p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our Kaggle community guidelines.</p>",
      "rawMarkdown": "**New to machine learning and data science?** No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\n**New to Kaggle?** Take a look at a few videos to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&v=GJBOMWpLpTQ). Publish and share your [models on Kaggle Models](https://www.kaggle.com/docs/models#publishing-a-model)!\n\n**Looking for a team?** Express your interest in joining a team through our [Team Up](https://www.kaggle.com/discussions/product-feedback/341195) feature .\n\n**Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our Kaggle community guidelines.",
      "votes": null
    },
    {
      "id": "2765557",
      "postDate": "04/21/2024 07:37:07",
      "content": "<p>Hi,I'm Begineer This dataset is huge How to run this huge dataset on local set my jupyter system start crashing</p>",
      "rawMarkdown": "Hi,I'm Begineer This dataset is huge How to run this huge dataset on local set my jupyter system start crashing",
      "votes": null
    },
    {
      "id": "2765660",
      "postDate": "04/21/2024 09:22:58",
      "content": "<p>Don't run on local computer<br>\nUse Google Colab because they provide free CPU and TPU runtimes</p>",
      "rawMarkdown": "Don't run on local computer\nUse Google Colab because they provide free CPU and TPU runtimes",
      "votes": null
    },
    {
      "id": "2772376",
      "postDate": "04/24/2024 17:29:33",
      "content": "<p>If you go to code and order by \"Most votes\" the top one shows you how to load portions of the data using duckdb</p>",
      "rawMarkdown": "If you go to code and order by \"Most votes\" the top one shows you how to load portions of the data using duckdb",
      "votes": null
    },
    {
      "id": "2772377",
      "postDate": "04/24/2024 17:32:49",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> Question on competition rules: Can we consult with online groups and people outside our team to get guidance and suggestions as long as we aren't privately sharing any of our code and they aren't writing any of our code? </p>",
      "rawMarkdown": "addisonhoward Question on competition rules: Can we consult with online groups and people outside our team to get guidance and suggestions as long as we aren't privately sharing any of our code and they aren't writing any of our code?",
      "votes": null
    },
    {
      "id": "2772454",
      "postDate": "04/24/2024 18:22:27",
      "content": "<p>Hi Nathaniel - those groups and people need to be publicly available. The former <em>may</em> be public (e.g. a question posed on Quora, StackOverflow, or a public Slack channel), the latter likely is not public as information gained from other people outside your team is not publicly available to others.</p>",
      "rawMarkdown": "Hi Nathaniel - those groups and people need to be publicly available. The former *may* be public (e.g. a question posed on Quora, StackOverflow, or a public Slack channel), the latter likely is not public as information gained from other people outside your team is not publicly available to others.",
      "votes": null
    },
    {
      "id": "2772460",
      "postDate": "04/24/2024 18:23:32",
      "content": "<p>A common example where this can be a rule violation: a class of students taking the same course all compete in the competition, but on different teams. But they share ideas about the competition offline without merging into a team within Kaggle.</p>",
      "rawMarkdown": "A common example where this can be a rule violation: a class of students taking the same course all compete in the competition, but on different teams. But they share ideas about the competition offline without merging into a team within Kaggle.",
      "votes": null
    },
    {
      "id": "2785039",
      "postDate": "04/30/2024 15:53:04",
      "content": "<p>Hi everybody. Right now, I can't open train file in my local computer because it's huge but will do on Colab. When I preview the train data on Kaggle, all bind values is equal to 0(Or I think so). How can I train my model If I don't have the relationship between molecules and their binds to proteins? Thanks a lot for help!!</p>\n<p>Edit: I think a minor portion of bind values is 1 but I am not sure. If It is, please let me know! Since I can't directly move through data itself and analyze it, I may be mistaking.</p>",
      "rawMarkdown": "Hi everybody. Right now, I can't open train file in my local computer because it's huge but will do on Colab. When I preview the train data on Kaggle, all bind values is equal to 0(Or I think so). How can I train my model If I don't have the relationship between molecules and their binds to proteins? Thanks a lot for help!!\n\nEdit: I think a minor portion of bind values is 1 but I am not sure. If It is, please let me know! Since I can't directly move through data itself and analyze it, I may be mistaking.",
      "votes": null
    },
    {
      "id": "2785157",
      "postDate": "04/30/2024 16:41:19",
      "content": "<p>There is a small number - about 1.5 million of the 295 million rows where binds = 1.  Check out many of the shared notebooks that use duckdb - they show how you can generate a dataset with binds = 0 and binds = 1.</p>",
      "rawMarkdown": "There is a small number - about 1.5 million of the 295 million rows where binds = 1.  Check out many of the shared notebooks that use duckdb - they show how you can generate a dataset with binds = 0 and binds = 1.",
      "votes": null
    },
    {
      "id": "2785168",
      "postDate": "04/30/2024 16:45:04",
      "content": "<p>Thanks Jimmy.</p>",
      "rawMarkdown": "Thanks Jimmy.",
      "votes": null
    },
    {
      "id": "2788304",
      "postDate": "05/02/2024 07:14:41",
      "content": "<p>hello community, i'm try to download the data using kaggle API but keep receiving the same error ' Traceback (most recent call last):<br>\n  File \"/usr/local/bin/kaggle\", line 8, in <br>\n    sys.exit(main())<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/cli.py\", line 54, in main<br>\n    out = args.func(**command_args)<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 1002, in competition_download_cli<br>\n    self.competition_download_files(competition, path, force,<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 965, in competition_download_files<br>\n    url = response.retries.history[0].redirect_location.split('?')[0]<br>\nIndexError: tuple index out of range'    , any help for this issue . thank you in advance </p>",
      "rawMarkdown": "hello community, i'm try to download the data using kaggle API but keep receiving the same error ' Traceback (most recent call last):\n  File \"/usr/local/bin/kaggle\", line 8, in <module>\n    sys.exit(main())\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/cli.py\", line 54, in main\n    out = args.func(**command_args)\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 1002, in competition_download_cli\n    self.competition_download_files(competition, path, force,\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 965, in competition_download_files\n    url = response.retries.history[0].redirect_location.split('?')[0]\nIndexError: tuple index out of range'    , any help for this issue . thank you in advance",
      "votes": null
    },
    {
      "id": "2792984",
      "postDate": "05/04/2024 13:23:17",
      "content": "<p>Hey everybody. I have 2 questions. Thanks for your answers.<br>\n1) Are all different building blocks crossed with each other? Which means there is no such order that can't create a molecule?<br>\n2) Is there any building block such that can be exist in more than just one column further down in data? Which means when molecules building blocks change places, they create a different molecule. Which also means the order between building blocks matters?</p>\n<p>Why I am asking these questions? Because when I multiply the numbers of different building blocks for **each **building block column, It is **not **equal to the number of different values in molecules column.</p>",
      "rawMarkdown": "Hey everybody. I have 2 questions. Thanks for your answers.\n1) Are all different building blocks crossed with each other? Which means there is no such order that can't create a molecule?\n2) Is there any building block such that can be exist in more than just one column further down in data? Which means when molecules building blocks change places, they create a different molecule. Which also means the order between building blocks matters?\n\nWhy I am asking these questions? Because when I multiply the numbers of different building blocks for **each **building block column, It is **not **equal to the number of different values in molecules column.",
      "votes": null
    },
    {
      "id": "2797312",
      "postDate": "05/06/2024 16:25:31",
      "content": "<p>I had a similar issue (not the same error, however) and it was because my kaggle.json was imported into the wrong location. </p>\n<pre><code>! ~/.kaggle\n! /content/.kaggle/kaggle.json ~/.kaggle/kaggle.json\n</code></pre>\n<p>Kind of a long shot since idk how your machine's configured, but this fixed it for me.</p>",
      "rawMarkdown": "I had a similar issue (not the same error, however) and it was because my kaggle.json was imported into the wrong location. \n```\n!mkdir ~/.kaggle\n!cp /content/.kaggle/kaggle.json ~/.kaggle/kaggle.json\n```\nKind of a long shot since idk how your machine's configured, but this fixed it for me.",
      "votes": null
    },
    {
      "id": "2798438",
      "postDate": "05/07/2024 08:23:57",
      "content": "<p>Kindly You can tell me What is this  Duchdb</p>",
      "rawMarkdown": "Kindly You can tell me What is this  Duchdb",
      "votes": null
    },
    {
      "id": "2802271",
      "postDate": "05/09/2024 00:32:32",
      "content": "<p>No. Changing place of molecules will certainly affect their structure and finally the target variable. So you should not change places of the molecules. </p>",
      "rawMarkdown": "No. Changing place of molecules will certainly affect their structure and finally the target variable. So you should not change places of the molecules.",
      "votes": null
    },
    {
      "id": "2802288",
      "postDate": "05/09/2024 01:16:50",
      "content": "<p>I tried different solutions including this method but still the same error , however, I read the data using paquet extension and it works . Thank you Travis</p>",
      "rawMarkdown": "I tried different solutions including this method but still the same error , however, I read the data using paquet extension and it works . Thank you Travis",
      "votes": null
    },
    {
      "id": "2889807",
      "postDate": "06/25/2024 17:15:51",
      "content": "<p>thank you for your sharing</p>",
      "rawMarkdown": "thank you for your sharing",
      "votes": null
    },
    {
      "id": "2894234",
      "postDate": "06/28/2024 10:47:13",
      "content": "<p>You can download the dataset and then train the same model in different chunked data each run</p>",
      "rawMarkdown": "You can download the dataset and then train the same model in different chunked data each run",
      "votes": null
    },
    {
      "id": "2902551",
      "postDate": "07/03/2024 10:36:13",
      "content": "<p>Hello Kagglers! Beginner here.</p>\n<p>I am trying to understand how the final scoring will work - I see that it is mentioned that a lot of unseen data will be added? But since you have to submit a .csv and not a code file, how will we get to run our models on this unseen data? Or have I misunderstood something completely? </p>\n<p>Hope you guys can help me out. Thanks!</p>",
      "rawMarkdown": "Hello Kagglers! Beginner here.\n\nI am trying to understand how the final scoring will work - I see that it is mentioned that a lot of unseen data will be added? But since you have to submit a .csv and not a code file, how will we get to run our models on this unseen data? Or have I misunderstood something completely? \n\nHope you guys can help me out. Thanks!",
      "votes": null
    },
    {
      "id": "2902871",
      "postDate": "07/03/2024 13:56:38",
      "content": "<p>Well! It is correct that you have a little of misunderstanding of Final Scoring. The .csv  File we  upload for prediction it will contrain predictions of all test data 100%   but Before the end of the Competion it will Use from the .csv file  about  for example 80%  predictions ( Originally it decided the Competetion Organizer) will be used to Caculate the score  from the our   file 100% predictions  which we upload  </p>",
      "rawMarkdown": "Well! It is correct that you have a little of misunderstanding of Final Scoring. The .csv  File we  upload for prediction it will contrain predictions of all test data 100%   but Before the end of the Competion it will Use from the .csv file  about  for example 80%  predictions ( Originally it decided the Competetion Organizer) will be used to Caculate the score  from the our   file 100% predictions  which we upload",
      "votes": null
    },
    {
      "id": "2902904",
      "postDate": "07/03/2024 14:20:16",
      "content": "<p>Okay, so they'll change the test.csv file at some point? When will this be? Do I have to be ready at some point then? </p>\n<p>Thank you for your answer :)</p>",
      "rawMarkdown": "Okay, so they'll change the test.csv file at some point? When will this be? Do I have to be ready at some point then? \n\nThank you for your answer :)",
      "votes": null
    },
    {
      "id": "2904634",
      "postDate": "07/04/2024 13:49:06",
      "content": "<p>Well there are two types of Competetion one is Code Competetion and other is not Code Competetion  the For the Code Competetion  We have to Write code  for training and prediction on test file and creating of this file .csv and When we submit our Note book it will run our code on Private Test data ('''means that it will not show to us some times  and  Some times an alias test data is shown to us  Which they will replace with original their test data sets  . ''') . In these types of Code competetion Test file s are Change only  but train data files remain same  Also we only submit Notebook in these competetion  </p>",
      "rawMarkdown": "Well there are two types of Competetion one is Code Competetion and other is not Code Competetion  the For the Code Competetion  We have to Write code  for training and prediction on test file and creating of this file .csv and When we submit our Note book it will run our code on Private Test data ('''means that it will not show to us some times  and  Some times an alias test data is shown to us  Which they will replace with original their test data sets  . ''') . In these types of Code competetion Test file s are Change only  but train data files remain same  Also we only submit Notebook in these competetion",
      "votes": null
    },
    {
      "id": "2904644",
      "postDate": "07/04/2024 13:51:49",
      "content": "<p>And second type of Competetion in which our test data file do  not change I have already explained above there is no changing of test  file Only  Selected Usually 80% prediction score are Calculated When Competetion Over Your WHOLE TEST FILE .CSV FILE SCORE AGAIN Calculated with your WHole prediction and then  Final score is assigned it will may change your position in leaderboard as well </p>",
      "rawMarkdown": "And second type of Competetion in which our test data file do  not change I have already explained above there is no changing of test  file Only  Selected Usually 80% prediction score are Calculated When Competetion Over Your WHOLE TEST FILE .CSV FILE SCORE AGAIN Calculated with your WHole prediction and then  Final score is assigned it will may change your position in leaderboard as well",
      "votes": null
    },
    {
      "id": "2904648",
      "postDate": "07/04/2024 13:55:54",
      "content": "<p>Thanks !!</p>",
      "rawMarkdown": "Thanks !!",
      "votes": null
    },
    {
      "id": "2909834",
      "postDate": "07/07/2024 11:07:31",
      "content": "<p>Ahh that makes sense! Thank you!</p>",
      "rawMarkdown": "Ahh that makes sense! Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2765557,
      "author_name": "mohammadsaquibdaiyan",
      "author_url": "",
      "post_date": "04/21/2024 07:37:07",
      "content": "<p>Hi,I'm Begineer This dataset is huge How to run this huge dataset on local set my jupyter system start crashing</p>",
      "votes": null,
      "replies": [
        {
          "id": 2765660,
          "author_name": "harsimrat11",
          "author_url": "",
          "post_date": "04/21/2024 09:22:58",
          "content": "<p>Don't run on local computer<br>\nUse Google Colab because they provide free CPU and TPU runtimes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2772376,
          "author_name": "nathanieldriggs",
          "author_url": "",
          "post_date": "04/24/2024 17:29:33",
          "content": "<p>If you go to code and order by \"Most votes\" the top one shows you how to load portions of the data using duckdb</p>",
          "votes": null,
          "replies": [
            {
              "id": 2798438,
              "author_name": "",
              "author_url": "",
              "post_date": "05/07/2024 08:23:57",
              "content": "<p>Kindly You can tell me What is this  Duchdb</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2904648,
              "author_name": "",
              "author_url": "",
              "post_date": "07/04/2024 13:55:54",
              "content": "<p>Thanks !!</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2894234,
          "author_name": "andreazedda91",
          "author_url": "",
          "post_date": "06/28/2024 10:47:13",
          "content": "<p>You can download the dataset and then train the same model in different chunked data each run</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2772377,
      "author_name": "nathanieldriggs",
      "author_url": "",
      "post_date": "04/24/2024 17:32:49",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> Question on competition rules: Can we consult with online groups and people outside our team to get guidance and suggestions as long as we aren't privately sharing any of our code and they aren't writing any of our code? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2772454,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "04/24/2024 18:22:27",
          "content": "<p>Hi Nathaniel - those groups and people need to be publicly available. The former <em>may</em> be public (e.g. a question posed on Quora, StackOverflow, or a public Slack channel), the latter likely is not public as information gained from other people outside your team is not publicly available to others.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2772460,
              "author_name": "addisonhoward",
              "author_url": "",
              "post_date": "04/24/2024 18:23:32",
              "content": "<p>A common example where this can be a rule violation: a class of students taking the same course all compete in the competition, but on different teams. But they share ideas about the competition offline without merging into a team within Kaggle.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2785039,
      "author_name": "erencil",
      "author_url": "",
      "post_date": "04/30/2024 15:53:04",
      "content": "<p>Hi everybody. Right now, I can't open train file in my local computer because it's huge but will do on Colab. When I preview the train data on Kaggle, all bind values is equal to 0(Or I think so). How can I train my model If I don't have the relationship between molecules and their binds to proteins? Thanks a lot for help!!</p>\n<p>Edit: I think a minor portion of bind values is 1 but I am not sure. If It is, please let me know! Since I can't directly move through data itself and analyze it, I may be mistaking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2785157,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "04/30/2024 16:41:19",
          "content": "<p>There is a small number - about 1.5 million of the 295 million rows where binds = 1.  Check out many of the shared notebooks that use duckdb - they show how you can generate a dataset with binds = 0 and binds = 1.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2785168,
              "author_name": "erencil",
              "author_url": "",
              "post_date": "04/30/2024 16:45:04",
              "content": "<p>Thanks Jimmy.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2788304,
      "author_name": "mouhsineelqesry",
      "author_url": "",
      "post_date": "05/02/2024 07:14:41",
      "content": "<p>hello community, i'm try to download the data using kaggle API but keep receiving the same error ' Traceback (most recent call last):<br>\n  File \"/usr/local/bin/kaggle\", line 8, in <br>\n    sys.exit(main())<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/cli.py\", line 54, in main<br>\n    out = args.func(**command_args)<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 1002, in competition_download_cli<br>\n    self.competition_download_files(competition, path, force,<br>\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 965, in competition_download_files<br>\n    url = response.retries.history[0].redirect_location.split('?')[0]<br>\nIndexError: tuple index out of range'    , any help for this issue . thank you in advance </p>",
      "votes": null,
      "replies": [
        {
          "id": 2797312,
          "author_name": "travislibre",
          "author_url": "",
          "post_date": "05/06/2024 16:25:31",
          "content": "<p>I had a similar issue (not the same error, however) and it was because my kaggle.json was imported into the wrong location. </p>\n<pre><code>! ~/.kaggle\n! /content/.kaggle/kaggle.json ~/.kaggle/kaggle.json\n</code></pre>\n<p>Kind of a long shot since idk how your machine's configured, but this fixed it for me.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2802288,
              "author_name": "mouhsineelqesry",
              "author_url": "",
              "post_date": "05/09/2024 01:16:50",
              "content": "<p>I tried different solutions including this method but still the same error , however, I read the data using paquet extension and it works . Thank you Travis</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2792984,
      "author_name": "erencil",
      "author_url": "",
      "post_date": "05/04/2024 13:23:17",
      "content": "<p>Hey everybody. I have 2 questions. Thanks for your answers.<br>\n1) Are all different building blocks crossed with each other? Which means there is no such order that can't create a molecule?<br>\n2) Is there any building block such that can be exist in more than just one column further down in data? Which means when molecules building blocks change places, they create a different molecule. Which also means the order between building blocks matters?</p>\n<p>Why I am asking these questions? Because when I multiply the numbers of different building blocks for **each **building block column, It is **not **equal to the number of different values in molecules column.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2802271,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "05/09/2024 00:32:32",
          "content": "<p>No. Changing place of molecules will certainly affect their structure and finally the target variable. So you should not change places of the molecules. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2889807,
      "author_name": "microgametezl",
      "author_url": "",
      "post_date": "06/25/2024 17:15:51",
      "content": "<p>thank you for your sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2902551,
      "author_name": "ebryhl",
      "author_url": "",
      "post_date": "07/03/2024 10:36:13",
      "content": "<p>Hello Kagglers! Beginner here.</p>\n<p>I am trying to understand how the final scoring will work - I see that it is mentioned that a lot of unseen data will be added? But since you have to submit a .csv and not a code file, how will we get to run our models on this unseen data? Or have I misunderstood something completely? </p>\n<p>Hope you guys can help me out. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2902871,
          "author_name": "",
          "author_url": "",
          "post_date": "07/03/2024 13:56:38",
          "content": "<p>Well! It is correct that you have a little of misunderstanding of Final Scoring. The .csv  File we  upload for prediction it will contrain predictions of all test data 100%   but Before the end of the Competion it will Use from the .csv file  about  for example 80%  predictions ( Originally it decided the Competetion Organizer) will be used to Caculate the score  from the our   file 100% predictions  which we upload  </p>",
          "votes": null,
          "replies": [
            {
              "id": 2902904,
              "author_name": "ebryhl",
              "author_url": "",
              "post_date": "07/03/2024 14:20:16",
              "content": "<p>Okay, so they'll change the test.csv file at some point? When will this be? Do I have to be ready at some point then? </p>\n<p>Thank you for your answer :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2904634,
                  "author_name": "",
                  "author_url": "",
                  "post_date": "07/04/2024 13:49:06",
                  "content": "<p>Well there are two types of Competetion one is Code Competetion and other is not Code Competetion  the For the Code Competetion  We have to Write code  for training and prediction on test file and creating of this file .csv and When we submit our Note book it will run our code on Private Test data ('''means that it will not show to us some times  and  Some times an alias test data is shown to us  Which they will replace with original their test data sets  . ''') . In these types of Code competetion Test file s are Change only  but train data files remain same  Also we only submit Notebook in these competetion  </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2904644,
                  "author_name": "",
                  "author_url": "",
                  "post_date": "07/04/2024 13:51:49",
                  "content": "<p>And second type of Competetion in which our test data file do  not change I have already explained above there is no changing of test  file Only  Selected Usually 80% prediction score are Calculated When Competetion Over Your WHOLE TEST FILE .CSV FILE SCORE AGAIN Calculated with your WHole prediction and then  Final score is assigned it will may change your position in leaderboard as well </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2909834,
                      "author_name": "ebryhl",
                      "author_url": "",
                      "post_date": "07/07/2024 11:07:31",
                      "content": "<p>Ahh that makes sense! Thank you!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2732053": "**New to machine learning and data science?** No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\n**New to Kaggle?** Take a look at a few videos to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&v=GJBOMWpLpTQ). Publish and share your [models on Kaggle Models](https://www.kaggle.com/docs/models#publishing-a-model)!\n\n**Looking for a team?** Express your interest in joining a team through our [Team Up](https://www.kaggle.com/discussions/product-feedback/341195) feature .\n\n**Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our Kaggle community guidelines.",
    "2765557": "Hi,I'm Begineer This dataset is huge How to run this huge dataset on local set my jupyter system start crashing",
    "2765660": "Don't run on local computer\nUse Google Colab because they provide free CPU and TPU runtimes",
    "2772376": "If you go to code and order by \"Most votes\" the top one shows you how to load portions of the data using duckdb",
    "2772377": "addisonhoward Question on competition rules: Can we consult with online groups and people outside our team to get guidance and suggestions as long as we aren't privately sharing any of our code and they aren't writing any of our code?",
    "2772454": "Hi Nathaniel - those groups and people need to be publicly available. The former *may* be public (e.g. a question posed on Quora, StackOverflow, or a public Slack channel), the latter likely is not public as information gained from other people outside your team is not publicly available to others.",
    "2772460": "A common example where this can be a rule violation: a class of students taking the same course all compete in the competition, but on different teams. But they share ideas about the competition offline without merging into a team within Kaggle.",
    "2785039": "Hi everybody. Right now, I can't open train file in my local computer because it's huge but will do on Colab. When I preview the train data on Kaggle, all bind values is equal to 0(Or I think so). How can I train my model If I don't have the relationship between molecules and their binds to proteins? Thanks a lot for help!!\n\nEdit: I think a minor portion of bind values is 1 but I am not sure. If It is, please let me know! Since I can't directly move through data itself and analyze it, I may be mistaking.",
    "2785157": "There is a small number - about 1.5 million of the 295 million rows where binds = 1.  Check out many of the shared notebooks that use duckdb - they show how you can generate a dataset with binds = 0 and binds = 1.",
    "2785168": "Thanks Jimmy.",
    "2788304": "hello community, i'm try to download the data using kaggle API but keep receiving the same error ' Traceback (most recent call last):\n  File \"/usr/local/bin/kaggle\", line 8, in <module>\n    sys.exit(main())\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/cli.py\", line 54, in main\n    out = args.func(**command_args)\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 1002, in competition_download_cli\n    self.competition_download_files(competition, path, force,\n  File \"/usr/local/lib/python3.10/dist-packages/kaggle/api/kaggle_api_extended.py\", line 965, in competition_download_files\n    url = response.retries.history[0].redirect_location.split('?')[0]\nIndexError: tuple index out of range'    , any help for this issue . thank you in advance",
    "2792984": "Hey everybody. I have 2 questions. Thanks for your answers.\n1) Are all different building blocks crossed with each other? Which means there is no such order that can't create a molecule?\n2) Is there any building block such that can be exist in more than just one column further down in data? Which means when molecules building blocks change places, they create a different molecule. Which also means the order between building blocks matters?\n\nWhy I am asking these questions? Because when I multiply the numbers of different building blocks for **each **building block column, It is **not **equal to the number of different values in molecules column.",
    "2797312": "I had a similar issue (not the same error, however) and it was because my kaggle.json was imported into the wrong location. \n```\n!mkdir ~/.kaggle\n!cp /content/.kaggle/kaggle.json ~/.kaggle/kaggle.json\n```\nKind of a long shot since idk how your machine's configured, but this fixed it for me.",
    "2798438": "Kindly You can tell me What is this  Duchdb",
    "2802271": "No. Changing place of molecules will certainly affect their structure and finally the target variable. So you should not change places of the molecules.",
    "2802288": "I tried different solutions including this method but still the same error , however, I read the data using paquet extension and it works . Thank you Travis",
    "2889807": "thank you for your sharing",
    "2894234": "You can download the dataset and then train the same model in different chunked data each run",
    "2902551": "Hello Kagglers! Beginner here.\n\nI am trying to understand how the final scoring will work - I see that it is mentioned that a lot of unseen data will be added? But since you have to submit a .csv and not a code file, how will we get to run our models on this unseen data? Or have I misunderstood something completely? \n\nHope you guys can help me out. Thanks!",
    "2902871": "Well! It is correct that you have a little of misunderstanding of Final Scoring. The .csv  File we  upload for prediction it will contrain predictions of all test data 100%   but Before the end of the Competion it will Use from the .csv file  about  for example 80%  predictions ( Originally it decided the Competetion Organizer) will be used to Caculate the score  from the our   file 100% predictions  which we upload",
    "2902904": "Okay, so they'll change the test.csv file at some point? When will this be? Do I have to be ready at some point then? \n\nThank you for your answer :)",
    "2904634": "Well there are two types of Competetion one is Code Competetion and other is not Code Competetion  the For the Code Competetion  We have to Write code  for training and prediction on test file and creating of this file .csv and When we submit our Note book it will run our code on Private Test data ('''means that it will not show to us some times  and  Some times an alias test data is shown to us  Which they will replace with original their test data sets  . ''') . In these types of Code competetion Test file s are Change only  but train data files remain same  Also we only submit Notebook in these competetion",
    "2904644": "And second type of Competetion in which our test data file do  not change I have already explained above there is no changing of test  file Only  Selected Usually 80% prediction score are Calculated When Competetion Over Your WHOLE TEST FILE .CSV FILE SCORE AGAIN Calculated with your WHole prediction and then  Final score is assigned it will may change your position in leaderboard as well",
    "2904648": "Thanks !!",
    "2909834": "Ahh that makes sense! Thank you!"
  },
  "source": "meta"
}