{
  "id": 384895,
  "title": "New to Machine Learning or Kaggle?",
  "url": "/competitions/asl-signs/discussion/384895",
  "author_name": "Ashley Chow",
  "post_date": "2023-02-10T04:44:16.263000",
  "votes": 10,
  "comment_count": 21,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>.</p>\n<p><strong>Remember:</strong> Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>",
  "messages": [
    {
      "id": 2137505,
      "postDate": "2023-02-10T04:44:16.263Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>.</p>\n<p><strong>Remember:</strong> Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle community guidelines</a>.</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&v=GJBOMWpLpTQ).\n\n**Remember:** Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).",
      "votes": 10
    },
    {
      "id": 2170100,
      "postDate": "2023-03-05T17:21:56.473Z",
      "content": "<p>Hi, pretty new here, first Kaggle competition, what infra are folks using in terms of training their models ? AWS EC2 instances, Amazon Sagemaker , or is there something more convenient to bootstrap faster ?</p>",
      "rawMarkdown": "Hi, pretty new here, first Kaggle competition, what infra are folks using in terms of training their models ? AWS EC2 instances, Amazon Sagemaker , or is there something more convenient to bootstrap faster ?",
      "votes": 1,
      "replies": [
        {
          "id": 2187836,
          "postDate": "2023-03-19T04:58:06.807Z",
          "content": "<p>Welcome to Kaggle and congratulations on your first competition!</p>\n<p>When it comes to infrastructure for training machine learning models, it really depends on the specific needs of the competition and the preferences of the individual competitor. AWS EC2 instances and Amazon SageMaker are both popular choices due to their scalability and flexibility, but they can also be more expensive.</p>\n<p>For competitors who are just starting out or have limited resources, there are more budget-friendly options such as using Kaggle's own cloud-based platform, Kaggle Notebooks, which provides free access to GPUs and TPUs. Other cloud providers, such as Google Cloud Platform and Microsoft Azure, also offer similar services.</p>\n<p>It's important to keep in mind that the infrastructure you use will depend on the specific requirements of the competition and the size of the dataset. In some cases, it may be possible to train models on your local machine, but this can be limited by hardware constraints and may not be feasible for larger datasets.</p>\n<p>The best approach is to assess the requirements of the competition and your own resources and choose the infrastructure that best suits your needs. Good luck with your competition!</p>",
          "rawMarkdown": "Welcome to Kaggle and congratulations on your first competition!\n\nWhen it comes to infrastructure for training machine learning models, it really depends on the specific needs of the competition and the preferences of the individual competitor. AWS EC2 instances and Amazon SageMaker are both popular choices due to their scalability and flexibility, but they can also be more expensive.\n\nFor competitors who are just starting out or have limited resources, there are more budget-friendly options such as using Kaggle's own cloud-based platform, Kaggle Notebooks, which provides free access to GPUs and TPUs. Other cloud providers, such as Google Cloud Platform and Microsoft Azure, also offer similar services.\n\nIt's important to keep in mind that the infrastructure you use will depend on the specific requirements of the competition and the size of the dataset. In some cases, it may be possible to train models on your local machine, but this can be limited by hardware constraints and may not be feasible for larger datasets.\n\nThe best approach is to assess the requirements of the competition and your own resources and choose the infrastructure that best suits your needs. Good luck with your competition!"
        }
      ]
    },
    {
      "id": 2167987,
      "postDate": "2023-03-03T19:53:18.747Z",
      "content": "<p>Any idea on what's the expected amount of hours to put in a competition to have any chances of winning? This seems very competitive.</p>",
      "rawMarkdown": "Any idea on what's the expected amount of hours to put in a competition to have any chances of winning? This seems very competitive.",
      "votes": 1,
      "replies": [
        {
          "id": 2187834,
          "postDate": "2023-03-19T04:57:10.373Z",
          "content": "<p>It's understandable to be curious about the amount of time it takes to be competitive in a Kaggle competition. The truth is, it can vary widely depending on the specific competition, your skill level, and how much time you're able to dedicate.</p>\n<p>Some competitions may require only a few hours of work, while others may require months of effort to be competitive. It's important to keep in mind that Kaggle attracts some of the best data scientists and machine learning practitioners in the world, so competition can be fierce.</p>\n<p>However, it's important to remember that competitions aren't just about winning. They can also be a great way to learn and improve your skills, as well as network with other professionals in the field. So even if you don't end up winning, you can still gain a lot from participating.</p>\n<p>Ultimately, the amount of time you put into a competition should depend on your personal goals and priorities. If you're just starting out, it may be more beneficial to focus on learning and building your skills, rather than trying to win a competition right away. With time and practice, you'll be able to compete at a higher level.</p>",
          "rawMarkdown": "It's understandable to be curious about the amount of time it takes to be competitive in a Kaggle competition. The truth is, it can vary widely depending on the specific competition, your skill level, and how much time you're able to dedicate.\n\nSome competitions may require only a few hours of work, while others may require months of effort to be competitive. It's important to keep in mind that Kaggle attracts some of the best data scientists and machine learning practitioners in the world, so competition can be fierce.\n\nHowever, it's important to remember that competitions aren't just about winning. They can also be a great way to learn and improve your skills, as well as network with other professionals in the field. So even if you don't end up winning, you can still gain a lot from participating.\n\nUltimately, the amount of time you put into a competition should depend on your personal goals and priorities. If you're just starting out, it may be more beneficial to focus on learning and building your skills, rather than trying to win a competition right away. With time and practice, you'll be able to compete at a higher level.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2159687,
      "postDate": "2023-02-26T00:01:13.200Z",
      "content": "<p>Hello , somewhat with the new kaggle competition submission approach.. technically i have a trained torch model. My understanding is that I will have to convert the torch model to tflite model to run predictions on the test set , but I'm not seeing the test.csv within the asl-sign folder.  </p>",
      "rawMarkdown": "Hello , somewhat with the new kaggle competition submission approach.. technically i have a trained torch model. My understanding is that I will have to convert the torch model to tflite model to run predictions on the test set , but I'm not seeing the test.csv within the asl-sign folder.  ",
      "votes": 1,
      "replies": [
        {
          "id": 2168264,
          "postDate": "2023-03-04T04:06:13.367Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2187830,
          "postDate": "2023-03-19T04:48:01.357Z",
          "content": "<p>The test set is hidden and will be used for final evaluation of the models and deciding who is the winners</p>",
          "rawMarkdown": "The test set is hidden and will be used for final evaluation of the models and deciding who is the winners"
        }
      ]
    },
    {
      "id": 2159683,
      "postDate": "2023-02-25T23:46:32.513Z",
      "content": "<p>Hi!<br>\nShould training be done using only submission notebook?<br>\nIs it acceptable to train the model locally and then upload it as a submission?</p>",
      "rawMarkdown": "Hi!\nShould training be done using only submission notebook?\nIs it acceptable to train the model locally and then upload it as a submission?",
      "votes": 1,
      "replies": [
        {
          "id": 2167046,
          "postDate": "2023-03-03T08:25:45.083Z",
          "content": "<p>Yeah, you can train locally</p>",
          "rawMarkdown": "Yeah, you can train locally"
        }
      ]
    },
    {
      "id": 2231991,
      "postDate": "2023-04-23T22:05:23.057Z",
      "content": "<p>Hi, I'm an aspiring data scientist and always looking for ways to improve my skills. I was wondering if anyone knows of any Kaggle competitions where the winning notebooks are public? I'd love to take a look and learn from the best.</p>",
      "rawMarkdown": "Hi, I'm an aspiring data scientist and always looking for ways to improve my skills. I was wondering if anyone knows of any Kaggle competitions where the winning notebooks are public? I'd love to take a look and learn from the best."
    },
    {
      "id": 2205389,
      "postDate": "2023-04-01T13:52:03.217Z",
      "content": "<p>how to choose which models are best suited?</p>",
      "rawMarkdown": "how to choose which models are best suited?"
    },
    {
      "id": 2203361,
      "postDate": "2023-03-30T18:42:24.533Z",
      "content": "<p>Hi there, I'm brand new to Kaggle, and i'm struggling with how to handle the data for this competition.<br>\nThe input is over 50 GB. but we're only given about 20 to work with. How can we handle data conversion? where do we store or make new datasets out of the cleaned and converted data(i made a wide version for example).<br>\nalso since inputs are read-only. How do we sample randomly from it since the parqet files are in subfolders and not in any sort of order?<br>\nany help would be greatly appreciated</p>",
      "rawMarkdown": "Hi there, I'm brand new to Kaggle, and i'm struggling with how to handle the data for this competition.\nThe input is over 50 GB. but we're only given about 20 to work with. How can we handle data conversion? where do we store or make new datasets out of the cleaned and converted data(i made a wide version for example).\nalso since inputs are read-only. How do we sample randomly from it since the parqet files are in subfolders and not in any sort of order?\nany help would be greatly appreciated"
    },
    {
      "id": 2187852,
      "postDate": "2023-03-19T05:06:40.217Z",
      "content": "<p>Hi, <br>\nThis may be a trivial question but if we are planning to only submit the tflite model, how should they replicate the preprocessing step?<br>\nfor example, I'm ignoring the z-axis and calculating some extra values, but now I'm not sure.<br>\nDo we have to strictly follow the <code>load_relevant_data_subset</code> function they have provided?<br>\nIf not should we include the changes there when submitting?<br>\nThanks for the help!</p>",
      "rawMarkdown": "Hi, \nThis may be a trivial question but if we are planning to only submit the tflite model, how should they replicate the preprocessing step?\nfor example, I'm ignoring the z-axis and calculating some extra values, but now I'm not sure.\nDo we have to strictly follow the `load_relevant_data_subset` function they have provided?\nIf not should we include the changes there when submitting?\nThanks for the help!",
      "replies": [
        {
          "id": 2190839,
          "postDate": "2023-03-21T14:24:13.103Z",
          "content": "<p>Hi,<br>\nI am in my first competition in Kaggle. I managed to submit a tflite model and get at least an ok-ish score. If you perform any kind of preprocessing for your model you have basically two options:</p>\n<ol>\n<li>Perform preprocessing inside your model already (loading directly the dataset with the <code>load_relevant_data_subset</code> or something similar).</li>\n<li>Perform preprocessing over all your dataset before training and then add that preprocessing step in your model before doing the submission.<br>\nI did option 2 with TensorFlow and my model for submission looks something like this (you can find a full example that I used as reference in this Notebook: <a href=\"https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language#Submit-Model):\" target=\"_blank\">https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language#Submit-Model):</a></li>\n</ol>\n<pre><code># Note that here we specify the size of a single input element\ninputs = layers.Input(shape=(CFG.ROWS_PER_FRAME, 3), name='inputs')\nx = Preprocessing()(inputs)\nx = tf.expand_dims(x, 0)\noutput = model(x)\n# explicitly name the final (identity) layer for the submission format\noutput = layers.Activation('linear', name='outputs')(output)\nmodel_with_preprocessing = tf.keras.Model(inputs=inputs, outputs=output)\n\n# Convert and save model\nconverter = tf.lite.TFLiteConverter.from_keras_model(model_with_preprocessing)\n# Do not apply optimization for faster output\n#converter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_model = converter.convert()\n\n# Save the model.\nwith open(CFG.OUTPUT_PATH + 'model.tflite', 'wb') as f:\n    f.write(tflite_model)\n\n# Compress model in .zip\nwith ZipFile(CFG.OUTPUT_PATH + 'submission.zip', 'w') as myzip:\n    myzip.write(CFG.OUTPUT_PATH + 'model.tflite')\n</code></pre>",
          "rawMarkdown": "Hi,\nI am in my first competition in Kaggle. I managed to submit a tflite model and get at least an ok-ish score. If you perform any kind of preprocessing for your model you have basically two options:\n1. Perform preprocessing inside your model already (loading directly the dataset with the `load_relevant_data_subset` or something similar).\n2. Perform preprocessing over all your dataset before training and then add that preprocessing step in your model before doing the submission.\nI did option 2 with TensorFlow and my model for submission looks something like this (you can find a full example that I used as reference in this Notebook: https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language#Submit-Model):\n```\n# Note that here we specify the size of a single input element\ninputs = layers.Input(shape=(CFG.ROWS_PER_FRAME, 3), name='inputs')\nx = Preprocessing()(inputs)\nx = tf.expand_dims(x, 0)\noutput = model(x)\n# explicitly name the final (identity) layer for the submission format\noutput = layers.Activation('linear', name='outputs')(output)\nmodel_with_preprocessing = tf.keras.Model(inputs=inputs, outputs=output)\n\n# Convert and save model\nconverter = tf.lite.TFLiteConverter.from_keras_model(model_with_preprocessing)\n# Do not apply optimization for faster output\n#converter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_model = converter.convert()\n\n# Save the model.\nwith open(CFG.OUTPUT_PATH + 'model.tflite', 'wb') as f:\n    f.write(tflite_model)\n    \n# Compress model in .zip\nwith ZipFile(CFG.OUTPUT_PATH + 'submission.zip', 'w') as myzip:\n    myzip.write(CFG.OUTPUT_PATH + 'model.tflite')\n```",
          "votes": 2,
          "replies": [
            {
              "id": 2204756,
              "postDate": "2023-03-31T22:44:07.333Z",
              "content": "<p>Many thanks :) <br>\nI'm working on the first option now, thought it is going to be a good learning experience </p>",
              "rawMarkdown": "Many thanks :) \nI'm working on the first option now, thought it is going to be a good learning experience "
            },
            {
              "id": 2220933,
              "postDate": "2023-04-13T19:32:59.990Z",
              "content": "<p>No problem! Best of lucks!</p>",
              "rawMarkdown": "No problem! Best of lucks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2176017,
      "postDate": "2023-03-10T10:24:21.480Z",
      "content": "<p>Hi… nice to meet you all…</p>",
      "rawMarkdown": "Hi... nice to meet you all..."
    },
    {
      "id": 2158896,
      "postDate": "2023-02-25T09:06:13.140Z",
      "content": "<p>hello ma'am. I am new in this field(machine learning).can you please explain me the project  \"Google - Isolated Sign Language Recognition\".</p>",
      "rawMarkdown": "hello ma'am. I am new in this field(machine learning).can you please explain me the project  \"Google - Isolated Sign Language Recognition\".\n",
      "replies": [
        {
          "id": 2179521,
          "postDate": "2023-03-13T08:26:43.467Z",
          "content": "<p>Hello me too ! That will be great.</p>",
          "rawMarkdown": "Hello me too ! That will be great.",
          "replies": [
            {
              "id": 2205955,
              "postDate": "2023-04-02T06:02:15.110Z",
              "content": "<p>hello me too !,  how I start the competition </p>",
              "rawMarkdown": "hello me too !,  how I start the competition "
            }
          ]
        }
      ]
    },
    {
      "id": 2171704,
      "postDate": "2023-03-07T03:05:08.993Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2170100,
      "author_name": "vinayp873",
      "author_url": "",
      "post_date": "2023-03-05T17:21:56.473000",
      "content": "<p>Hi, pretty new here, first Kaggle competition, what infra are folks using in terms of training their models ? AWS EC2 instances, Amazon Sagemaker , or is there something more convenient to bootstrap faster ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2187836,
          "author_name": "Siddharth Kumar",
          "author_url": "",
          "post_date": "2023-03-19T04:58:06.807000",
          "content": "<p>Welcome to Kaggle and congratulations on your first competition!</p>\n<p>When it comes to infrastructure for training machine learning models, it really depends on the specific needs of the competition and the preferences of the individual competitor. AWS EC2 instances and Amazon SageMaker are both popular choices due to their scalability and flexibility, but they can also be more expensive.</p>\n<p>For competitors who are just starting out or have limited resources, there are more budget-friendly options such as using Kaggle's own cloud-based platform, Kaggle Notebooks, which provides free access to GPUs and TPUs. Other cloud providers, such as Google Cloud Platform and Microsoft Azure, also offer similar services.</p>\n<p>It's important to keep in mind that the infrastructure you use will depend on the specific requirements of the competition and the size of the dataset. In some cases, it may be possible to train models on your local machine, but this can be limited by hardware constraints and may not be feasible for larger datasets.</p>\n<p>The best approach is to assess the requirements of the competition and your own resources and choose the infrastructure that best suits your needs. Good luck with your competition!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2167987,
      "author_name": "salah488",
      "author_url": "",
      "post_date": "2023-03-03T19:53:18.747000",
      "content": "<p>Any idea on what's the expected amount of hours to put in a competition to have any chances of winning? This seems very competitive.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2187834,
          "author_name": "Siddharth Kumar",
          "author_url": "",
          "post_date": "2023-03-19T04:57:10.373000",
          "content": "<p>It's understandable to be curious about the amount of time it takes to be competitive in a Kaggle competition. The truth is, it can vary widely depending on the specific competition, your skill level, and how much time you're able to dedicate.</p>\n<p>Some competitions may require only a few hours of work, while others may require months of effort to be competitive. It's important to keep in mind that Kaggle attracts some of the best data scientists and machine learning practitioners in the world, so competition can be fierce.</p>\n<p>However, it's important to remember that competitions aren't just about winning. They can also be a great way to learn and improve your skills, as well as network with other professionals in the field. So even if you don't end up winning, you can still gain a lot from participating.</p>\n<p>Ultimately, the amount of time you put into a competition should depend on your personal goals and priorities. If you're just starting out, it may be more beneficial to focus on learning and building your skills, rather than trying to win a competition right away. With time and practice, you'll be able to compete at a higher level.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2159687,
      "author_name": "Siaw-Darko Emmanuel",
      "author_url": "",
      "post_date": "2023-02-26T00:01:13.200000",
      "content": "<p>Hello , somewhat with the new kaggle competition submission approach.. technically i have a trained torch model. My understanding is that I will have to convert the torch model to tflite model to run predictions on the test set , but I'm not seeing the test.csv within the asl-sign folder.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2168264,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-04T04:06:13.367000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2187830,
          "author_name": "Rasha Salim",
          "author_url": "",
          "post_date": "2023-03-19T04:48:01.357000",
          "content": "<p>The test set is hidden and will be used for final evaluation of the models and deciding who is the winners</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2159683,
      "author_name": "Kamil",
      "author_url": "",
      "post_date": "2023-02-25T23:46:32.513000",
      "content": "<p>Hi!<br>\nShould training be done using only submission notebook?<br>\nIs it acceptable to train the model locally and then upload it as a submission?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2167046,
          "author_name": "n0n1ck",
          "author_url": "",
          "post_date": "2023-03-03T08:25:45.083000",
          "content": "<p>Yeah, you can train locally</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2231991,
      "author_name": "Arnau Vilella Piqué",
      "author_url": "",
      "post_date": "2023-04-23T22:05:23.057000",
      "content": "<p>Hi, I'm an aspiring data scientist and always looking for ways to improve my skills. I was wondering if anyone knows of any Kaggle competitions where the winning notebooks are public? I'd love to take a look and learn from the best.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2205389,
      "author_name": "erlan chodoev",
      "author_url": "",
      "post_date": "2023-04-01T13:52:03.217000",
      "content": "<p>how to choose which models are best suited?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2203361,
      "author_name": "Avi Feygin",
      "author_url": "",
      "post_date": "2023-03-30T18:42:24.533000",
      "content": "<p>Hi there, I'm brand new to Kaggle, and i'm struggling with how to handle the data for this competition.<br>\nThe input is over 50 GB. but we're only given about 20 to work with. How can we handle data conversion? where do we store or make new datasets out of the cleaned and converted data(i made a wide version for example).<br>\nalso since inputs are read-only. How do we sample randomly from it since the parqet files are in subfolders and not in any sort of order?<br>\nany help would be greatly appreciated</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2187852,
      "author_name": "Rasha Salim",
      "author_url": "",
      "post_date": "2023-03-19T05:06:40.217000",
      "content": "<p>Hi, <br>\nThis may be a trivial question but if we are planning to only submit the tflite model, how should they replicate the preprocessing step?<br>\nfor example, I'm ignoring the z-axis and calculating some extra values, but now I'm not sure.<br>\nDo we have to strictly follow the <code>load_relevant_data_subset</code> function they have provided?<br>\nIf not should we include the changes there when submitting?<br>\nThanks for the help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2190839,
          "author_name": "Miguel Trasobares Baselga",
          "author_url": "",
          "post_date": "2023-03-21T14:24:13.103000",
          "content": "<p>Hi,<br>\nI am in my first competition in Kaggle. I managed to submit a tflite model and get at least an ok-ish score. If you perform any kind of preprocessing for your model you have basically two options:</p>\n<ol>\n<li>Perform preprocessing inside your model already (loading directly the dataset with the <code>load_relevant_data_subset</code> or something similar).</li>\n<li>Perform preprocessing over all your dataset before training and then add that preprocessing step in your model before doing the submission.<br>\nI did option 2 with TensorFlow and my model for submission looks something like this (you can find a full example that I used as reference in this Notebook: <a href=\"https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language#Submit-Model):\" target=\"_blank\">https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language#Submit-Model):</a></li>\n</ol>\n<pre><code># Note that here we specify the size of a single input element\ninputs = layers.Input(shape=(CFG.ROWS_PER_FRAME, 3), name='inputs')\nx = Preprocessing()(inputs)\nx = tf.expand_dims(x, 0)\noutput = model(x)\n# explicitly name the final (identity) layer for the submission format\noutput = layers.Activation('linear', name='outputs')(output)\nmodel_with_preprocessing = tf.keras.Model(inputs=inputs, outputs=output)\n\n# Convert and save model\nconverter = tf.lite.TFLiteConverter.from_keras_model(model_with_preprocessing)\n# Do not apply optimization for faster output\n#converter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_model = converter.convert()\n\n# Save the model.\nwith open(CFG.OUTPUT_PATH + 'model.tflite', 'wb') as f:\n    f.write(tflite_model)\n\n# Compress model in .zip\nwith ZipFile(CFG.OUTPUT_PATH + 'submission.zip', 'w') as myzip:\n    myzip.write(CFG.OUTPUT_PATH + 'model.tflite')\n</code></pre>",
          "votes": 2,
          "replies": [
            {
              "id": 2204756,
              "author_name": "Rasha Salim",
              "author_url": "",
              "post_date": "2023-03-31T22:44:07.333000",
              "content": "<p>Many thanks :) <br>\nI'm working on the first option now, thought it is going to be a good learning experience </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2220933,
              "author_name": "Miguel Trasobares Baselga",
              "author_url": "",
              "post_date": "2023-04-13T19:32:59.990000",
              "content": "<p>No problem! Best of lucks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2176017,
      "author_name": "Samuel Robert Ardi Nugraha",
      "author_url": "",
      "post_date": "2023-03-10T10:24:21.480000",
      "content": "<p>Hi… nice to meet you all…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2158896,
      "author_name": "DEBAD909",
      "author_url": "",
      "post_date": "2023-02-25T09:06:13.140000",
      "content": "<p>hello ma'am. I am new in this field(machine learning).can you please explain me the project  \"Google - Isolated Sign Language Recognition\".</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2179521,
          "author_name": "wael ben dahou",
          "author_url": "",
          "post_date": "2023-03-13T08:26:43.467000",
          "content": "<p>Hello me too ! That will be great.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2205955,
              "author_name": "Uttam Kumar",
              "author_url": "",
              "post_date": "2023-04-02T06:02:15.110000",
              "content": "<p>hello me too !,  how I start the competition </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2171704,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-07T03:05:08.993000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2137505": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&v=GJBOMWpLpTQ).\n\n**Remember:** Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).",
    "2170100": "Hi, pretty new here, first Kaggle competition, what infra are folks using in terms of training their models ? AWS EC2 instances, Amazon Sagemaker , or is there something more convenient to bootstrap faster ?",
    "2167987": "Any idea on what's the expected amount of hours to put in a competition to have any chances of winning? This seems very competitive.",
    "2159687": "Hello , somewhat with the new kaggle competition submission approach.. technically i have a trained torch model. My understanding is that I will have to convert the torch model to tflite model to run predictions on the test set , but I'm not seeing the test.csv within the asl-sign folder.  ",
    "2159683": "Hi!\nShould training be done using only submission notebook?\nIs it acceptable to train the model locally and then upload it as a submission?",
    "2231991": "Hi, I'm an aspiring data scientist and always looking for ways to improve my skills. I was wondering if anyone knows of any Kaggle competitions where the winning notebooks are public? I'd love to take a look and learn from the best.",
    "2205389": "how to choose which models are best suited?",
    "2203361": "Hi there, I'm brand new to Kaggle, and i'm struggling with how to handle the data for this competition.\nThe input is over 50 GB. but we're only given about 20 to work with. How can we handle data conversion? where do we store or make new datasets out of the cleaned and converted data(i made a wide version for example).\nalso since inputs are read-only. How do we sample randomly from it since the parqet files are in subfolders and not in any sort of order?\nany help would be greatly appreciated",
    "2187852": "Hi, \nThis may be a trivial question but if we are planning to only submit the tflite model, how should they replicate the preprocessing step?\nfor example, I'm ignoring the z-axis and calculating some extra values, but now I'm not sure.\nDo we have to strictly follow the `load_relevant_data_subset` function they have provided?\nIf not should we include the changes there when submitting?\nThanks for the help!",
    "2176017": "Hi... nice to meet you all...",
    "2158896": "hello ma'am. I am new in this field(machine learning).can you please explain me the project  \"Google - Isolated Sign Language Recognition\".\n",
    "2171704": ""
  }
}