{
  "id": 101221,
  "title": "[current LB: 0.435] Kicking the tires on Google AutoML Vision ",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101221",
  "author_name": "",
  "post_date": "2019-07-24T05:30:14.244045700Z",
  "votes": 40,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I’m a long-time competition lurker who's finally mustered the courage to join a competition 😅. I picked this competition because I find the topic interesting (if I was going to college today I’d probably pick biology). And also as a chance to try out <a href=\"https://cloud.google.com/vision/automl/docs/\">Google AutoML Vision</a>. </p>\n\n<p>Kaggle is part of Google now, which gives me exposure to some of the technologies that are generating excitement inside Google. AutoML is a tool with a lot of buzz. </p>\n\n<p>I hadn’t paid much attention to the Automated Machine Learning Tools (AMLTs) until the KaggleDays competition in San Francisco this past April. Two AMLTs, Google AutoML and H20, ended up participating in the hackathon and getting top 10 performances (here’s the <a href=\"https://ai.googleblog.com/2019/05/an-end-to-end-automl-solution-for.html\">Google AI writeup</a> of their performance). While I believe it requires human creativity to do problem setup and domain-specific feature engineering, KaggleDays left me wondering whether humans need to be doing generic feature engineering, architecture selection, picking activation functions and setting learning rates. </p>\n\n<p>I plan to share details of my experiments and experiences with Google AutoML Vision in this thread. I’m also open to suggestions from others in the community for things I should try. But bear with me if I take some-time to respond: like many other Kagglers I have a busy day job.</p>",
  "messages": [
    {
      "id": "583151",
      "postDate": "07/24/2019 05:30:14",
      "content": "<p>I’m a long-time competition lurker who's finally mustered the courage to join a competition 😅. I picked this competition because I find the topic interesting (if I was going to college today I’d probably pick biology). And also as a chance to try out <a href=\"https://cloud.google.com/vision/automl/docs/\">Google AutoML Vision</a>. </p>\n\n<p>Kaggle is part of Google now, which gives me exposure to some of the technologies that are generating excitement inside Google. AutoML is a tool with a lot of buzz. </p>\n\n<p>I hadn’t paid much attention to the Automated Machine Learning Tools (AMLTs) until the KaggleDays competition in San Francisco this past April. Two AMLTs, Google AutoML and H20, ended up participating in the hackathon and getting top 10 performances (here’s the <a href=\"https://ai.googleblog.com/2019/05/an-end-to-end-automl-solution-for.html\">Google AI writeup</a> of their performance). While I believe it requires human creativity to do problem setup and domain-specific feature engineering, KaggleDays left me wondering whether humans need to be doing generic feature engineering, architecture selection, picking activation functions and setting learning rates. </p>\n\n<p>I plan to share details of my experiments and experiences with Google AutoML Vision in this thread. I’m also open to suggestions from others in the community for things I should try. But bear with me if I take some-time to respond: like many other Kagglers I have a busy day job.</p>",
      "rawMarkdown": "I’m a long-time competition lurker who's finally mustered the courage to join a competition 😅. I picked this competition because I find the topic interesting (if I was going to college today I’d probably pick biology). And also as a chance to try out [Google AutoML Vision](https://cloud.google.com/vision/automl/docs/). \n\nKaggle is part of Google now, which gives me exposure to some of the technologies that are generating excitement inside Google. AutoML is a tool with a lot of buzz. \n\nI hadn’t paid much attention to the Automated Machine Learning Tools (AMLTs) until the KaggleDays competition in San Francisco this past April. Two AMLTs, Google AutoML and H20, ended up participating in the hackathon and getting top 10 performances (here’s the [Google AI writeup](https://ai.googleblog.com/2019/05/an-end-to-end-automl-solution-for.html) of their performance). While I believe it requires human creativity to do problem setup and domain-specific feature engineering, KaggleDays left me wondering whether humans need to be doing generic feature engineering, architecture selection, picking activation functions and setting learning rates. \n\nI plan to share details of my experiments and experiences with Google AutoML Vision in this thread. I’m also open to suggestions from others in the community for things I should try. But bear with me if I take some-time to respond: like many other Kagglers I have a busy day job.",
      "votes": null
    },
    {
      "id": "583156",
      "postDate": "07/24/2019 05:34:52",
      "content": "<h2>First Few Experiments</h2>\n\n<p>I started with the images created by <a href=\"/xhlulu\">@xhlulu</a>’s <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">kernel</a> (thank you <a href=\"/xhlulu\">@xhlulu</a>!). In that kernel, xhlulu@ converts the data to 224x224 RGB jpegs. </p>\n\n<p>By default, Google AutoML randomly splits between TRAIN, VALIDATION and TEST. And the predicted outputs labels and confidence scores for each image (I define a threshold and it outputs all labels and confidences above that confidence threshold). </p>\n\n<p>I made predictions for both sites and then picked the label that had a higher confidence score across both sites. Below is a table of my scores. (Note: AUC is the accuracy metric that Google AutoML Vision reports. Haven’t put the time into figuring out what it means but I assume it’s probably treating each class as binary on/off and then aggregating up.)</p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Public Leaderboard Score\n   \n  \n  \n   24\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.23\n   \n   0.101\n   \n  \n  \n   48\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.42\n   \n   \n   \n  \n  \n   72\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.5\n   \n   \n   \n  \n  \n   96\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.56\n   \n   0.199\n   \n  \n\n\n<p>I then made small modifications xhlulu@’s code to create 512x512 RGB pngs to test the impact of higher resolution images.</p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Leaderboard Score\n   \n  \n  \n   24\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.32\n   \n   \n   \n  \n  \n   48\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.31\n   \n   \n   \n  \n  \n   72\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.54\n   \n   0.254\n   \n  \n  \n   96\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.64\n   \n   0.272\n   \n  \n  \n   120\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.68\n   \n   0.292\n   \n  \n\n\n<p>0.302 is my best score. It’s a result of a poorly thought out experiment, so I’m surprised it’s my best score. I trained a model for 144 hours that included the control images. But after inspecting the test-set predictions I realized my model was often picking labels &gt; 1108 (which are control labels that don’t appear in the test set). I ended up filtering the test-set predictions to go to the next most confident label when a label &gt;1108 was predicted. </p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Leaderboard Score\n   \n  \n  \n   24\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.34\n   \n   0.175\n   \n  \n  \n   48\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.54\n   \n   \n   \n  \n  \n   72\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.63\n   \n   \n   \n  \n  \n   96\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.68\n   \n   0.283\n   \n  \n  \n   120\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.71\n   \n   0.294\n   \n  \n  \n   144\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.72\n   \n   0.302\n   \n  \n\n\n<p>Google AutoML Vision allows me to define my own TRAIN, VALIDATION and TEST splits. For my next experiments I’m going to start using this feature. I’m going to define my own splits to:\n1. Include the control data in TRAIN but not VALIDATION or TEST (to allow the model to use the control data for training but to discourage the model from predicting labels &gt;1108 on the competition test set)\n2. Split by experiment rather than randomly </p>",
      "rawMarkdown": "## First Few Experiments\n\nI started with the images created by @xhlulu’s [kernel](https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images) (thank you @xhlulu!). In that kernel, xhlulu@ converts the data to 224x224 RGB jpegs. \n\nBy default, Google AutoML randomly splits between TRAIN, VALIDATION and TEST. And the predicted outputs labels and confidence scores for each image (I define a threshold and it outputs all labels and confidences above that confidence threshold). \n\nI made predictions for both sites and then picked the label that had a higher confidence score across both sites. Below is a table of my scores. (Note: AUC is the accuracy metric that Google AutoML Vision reports. Haven’t put the time into figuring out what it means but I assume it’s probably treating each class as binary on/off and then aggregating up.)\n\n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Public Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.23\n   </td>\n   <td>0.101\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.42\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.5\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.56\n   </td>\n   <td>0.199\n   </td>\n  </tr>\n</tbody></table>\n\n\nI then made small modifications xhlulu@’s code to create 512x512 RGB pngs to test the impact of higher resolution images.\n\n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.32\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.31\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.54\n   </td>\n   <td>0.254\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.64\n   </td>\n   <td>0.272\n   </td>\n  </tr>\n  <tr>\n   <td>120\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.68\n   </td>\n   <td>0.292\n   </td>\n  </tr>\n</tbody></table>\n\n0.302 is my best score. It’s a result of a poorly thought out experiment, so I’m surprised it’s my best score. I trained a model for 144 hours that included the control images. But after inspecting the test-set predictions I realized my model was often picking labels &gt; 1108 (which are control labels that don’t appear in the test set). I ended up filtering the test-set predictions to go to the next most confident label when a label &gt;1108 was predicted. \n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.34\n   </td>\n   <td>0.175\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.54\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.63\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.68\n   </td>\n   <td>0.283\n   </td>\n  </tr>\n  <tr>\n   <td>120\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.71\n   </td>\n   <td>0.294\n   </td>\n  </tr>\n  <tr>\n   <td>144\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.72\n   </td>\n   <td>0.302\n   </td>\n  </tr>\n</tbody></table>\n\nGoogle AutoML Vision allows me to define my own TRAIN, VALIDATION and TEST splits. For my next experiments I’m going to start using this feature. I’m going to define my own splits to:\n1. Include the control data in TRAIN but not VALIDATION or TEST (to allow the model to use the control data for training but to discourage the model from predicting labels &gt;1108 on the competition test set)\n2. Split by experiment rather than randomly",
      "votes": null
    },
    {
      "id": "583158",
      "postDate": "07/24/2019 05:40:40",
      "content": "<h2>Experience with Google AutoML Vision</h2>\n\n<p>I’m getting pretty decent results considering how naive my models are. I’ve found Google AutoML easy in many ways. To get a basic model all I need to do is:\n1.  upload the RGB images to Google Cloud Storage\n2.  upload a CSV file with a pointer to the GCS bucket and the target label</p>\n\n<p>There are a bunch of limitations with the products. My biggest issues so far have been:\n1. inability to add other metadata (e.g. would like to be able to add metadata on controls, plate id and position on plate). I can probably get around this using a post-processing step but it’d be nice to be able to add this metadata into the single model.\n2. there isn’t a feature that allows batch prediction on a large number of images. I work around this by hitting the prediction API 19897 times to generate my submission file. \n3. the model evaluation page is not very helpful for debugging model performance. <br>\n4. I had to use a lot of hours of training to get good results (you can see that the models keep improving with additional training time). This means I have to wait a long time and spend a lot of money.\n5. can’t change the loss function. Based on the forums, it seems like the loss function might be an important setting for this competition. </p>\n\n<p>There are some more minor frustrations that I had to workaround, which I’m happy to share if others are planning to try it out and want to learn from some of the frictions I encountered.</p>",
      "rawMarkdown": "## Experience with Google AutoML Vision\n\nI’m getting pretty decent results considering how naive my models are. I’ve found Google AutoML easy in many ways. To get a basic model all I need to do is:\n1.  upload the RGB images to Google Cloud Storage\n2.  upload a CSV file with a pointer to the GCS bucket and the target label\n\nThere are a bunch of limitations with the products. My biggest issues so far have been:\n1. inability to add other metadata (e.g. would like to be able to add metadata on controls, plate id and position on plate). I can probably get around this using a post-processing step but it’d be nice to be able to add this metadata into the single model.\n2. there isn’t a feature that allows batch prediction on a large number of images. I work around this by hitting the prediction API 19897 times to generate my submission file. \n3. the model evaluation page is not very helpful for debugging model performance.  \n4. I had to use a lot of hours of training to get good results (you can see that the models keep improving with additional training time). This means I have to wait a long time and spend a lot of money.\n5. can’t change the loss function. Based on the forums, it seems like the loss function might be an important setting for this competition. \n\nThere are some more minor frustrations that I had to workaround, which I’m happy to share if others are planning to try it out and want to learn from some of the frictions I encountered.",
      "votes": null
    },
    {
      "id": "583248",
      "postDate": "07/24/2019 08:09:33",
      "content": "<p>Hi Anthony, Nice to meet you!!! You are an inspiration ;)</p>",
      "rawMarkdown": "Hi Anthony, Nice to meet you!!! You are an inspiration ;)",
      "votes": null
    },
    {
      "id": "583285",
      "postDate": "07/24/2019 09:28:11",
      "content": "<p>Thank for sharing. Looking at your training hours, I think I should have more patience with my model.</p>",
      "rawMarkdown": "Thank for sharing. Looking at your training hours, I think I should have more patience with my model.",
      "votes": null
    },
    {
      "id": "583489",
      "postDate": "07/24/2019 14:37:39",
      "content": "<p><a href=\"/lkhphuc\">@lkhphuc</a>, unlikely you need anywhere near as many training hours if you're not using AutoML. The upside of AutoML is that you don't have to set any hyperparameters and still get a decent model. One of the drawbacks is it's so compute intensive. </p>",
      "rawMarkdown": "lkhphuc, unlikely you need anywhere near as many training hours if you're not using AutoML. The upside of AutoML is that you don't have to set any hyperparameters and still get a decent model. One of the drawbacks is it's so compute intensive.",
      "votes": null
    },
    {
      "id": "583525",
      "postDate": "07/24/2019 15:22:18",
      "content": "<p>Those results are pretty promising, considering that there was no architecture tuning whatsoever. The massive improvement in accuracy when you change the input from 224px jpegs to 512px pngs also seems to indicate that the compression format and size of the images plays an important factor in the accuracy.</p>\n\n<p>This is surprising in some way, since most of the classical architectures are designed to receive as input image sizes between 200 and 300, and greater image sizes usually requires training a filter that downsample the image into a size accepted by the network, or else it could result in OOM issues.</p>\n\n<p>If you still have the 512px png dataset laying around, could you possibly share it as a public dataset? That would save some of us spending hours generating the dataset locally ;)</p>",
      "rawMarkdown": "Those results are pretty promising, considering that there was no architecture tuning whatsoever. The massive improvement in accuracy when you change the input from 224px jpegs to 512px pngs also seems to indicate that the compression format and size of the images plays an important factor in the accuracy.\n\nThis is surprising in some way, since most of the classical architectures are designed to receive as input image sizes between 200 and 300, and greater image sizes usually requires training a filter that downsample the image into a size accepted by the network, or else it could result in OOM issues.\n\nIf you still have the 512px png dataset laying around, could you possibly share it as a public dataset? That would save some of us spending hours generating the dataset locally ;)",
      "votes": null
    },
    {
      "id": "583537",
      "postDate": "07/24/2019 15:33:10",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a>, I actually regret converting to png. It would have been a purer comparison if I'd also used jpeg. </p>\n\n<p>I tried running the 512px conversion in a kernel but I didn't have quite enough compute time. Let me find out how I can share the 512px dataset with everyone. </p>",
      "rawMarkdown": "xhlulu, I actually regret converting to png. It would have been a purer comparison if I'd also used jpeg. \n\nI tried running the 512px conversion in a kernel but I didn't have quite enough compute time. Let me find out how I can share the 512px dataset with everyone.",
      "votes": null
    },
    {
      "id": "583539",
      "postDate": "07/24/2019 15:34:30",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "583711",
      "postDate": "07/24/2019 21:29:38",
      "content": "<p>Very interesting to see your code. Can Google AutoML be run on kernels (of course, there is a running time limitation of 9h, but still)? Maybe you can publish on github? </p>\n\n<p>Does it give architecture and hyper parameters recommendations that can be used separately afterwards?</p>",
      "rawMarkdown": "Very interesting to see your code. Can Google AutoML be run on kernels (of course, there is a running time limitation of 9h, but still)? Maybe you can publish on github? \n\nDoes it give architecture and hyper parameters recommendations that can be used separately afterwards?",
      "votes": null
    },
    {
      "id": "583759",
      "postDate": "07/25/2019 00:10:31",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a>, unfortunately can't share the dataset through the datasets platform. That'd allow users to access the dataset without having to accept the competitions rules. </p>",
      "rawMarkdown": "xhlulu, unfortunately can't share the dataset through the datasets platform. That'd allow users to access the dataset without having to accept the competitions rules.",
      "votes": null
    },
    {
      "id": "585543",
      "postDate": "07/27/2019 15:55:45",
      "content": "<p>Thank you for sharing this awesome tool! May I know what's your GPU resource?</p>",
      "rawMarkdown": "Thank you for sharing this awesome tool! May I know what's your GPU resource?",
      "votes": null
    },
    {
      "id": "589256",
      "postDate": "07/31/2019 17:12:11",
      "content": "<p><a href=\"/zaharch\">@zaharch</a>, sorry for the delay. I wanted to see if AutoML could be run out of Kernels. Turns out it can:\n<a href=\"https://www.kaggle.com/antgoldbloom/training-using-google-automl/\">https://www.kaggle.com/antgoldbloom/training-using-google-automl/</a></p>\n\n<p>The 9 hour limit for kernels is not an issue because my kernel kicks off a 24 hour training job and then stops. </p>\n\n<p>My next step is to share my inference code which sends the test images to the trained model and generates predicted labels. </p>\n\n<p>AutoML doesn't give architecture or hyper parameter recommendations. It just creates an end point that you can send images to and get back predicted labels (and the model's confidence in that label).  </p>",
      "rawMarkdown": "zaharch, sorry for the delay. I wanted to see if AutoML could be run out of Kernels. Turns out it can:\nhttps://www.kaggle.com/antgoldbloom/training-using-google-automl/\n\nThe 9 hour limit for kernels is not an issue because my kernel kicks off a 24 hour training job and then stops. \n\nMy next step is to share my inference code which sends the test images to the trained model and generates predicted labels. \n\nAutoML doesn't give architecture or hyper parameter recommendations. It just creates an end point that you can send images to and get back predicted labels (and the model's confidence in that label).",
      "votes": null
    },
    {
      "id": "589258",
      "postDate": "07/31/2019 17:13:40",
      "content": "<p><a href=\"/ttylacm\">@ttylacm</a>, don't know what hardware Google AutoML is using under the hood. It's invisible to the user. Good chance it's using TPUs though.  </p>",
      "rawMarkdown": "ttylacm, don't know what hardware Google AutoML is using under the hood. It's invisible to the user. Good chance it's using TPUs though.",
      "votes": null
    },
    {
      "id": "589262",
      "postDate": "07/31/2019 17:14:58",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks! Looks like you're a few places ahead of me on the leaderboard. Watch out, I have a few new ideas to try 😏 </p>",
      "rawMarkdown": "ratthachat thanks! Looks like you're a few places ahead of me on the leaderboard. Watch out, I have a few new ideas to try 😏",
      "votes": null
    },
    {
      "id": "589452",
      "postDate": "07/31/2019 23:04:13",
      "content": "<p><a href=\"/antgoldbloom\">@antgoldbloom</a> That’s interesting Anthony! I just received GCP credits from the competition and just finished setting the (preemptible) GPU. So I can also test more fancy methods next week.</p>",
      "rawMarkdown": "antgoldbloom That’s interesting Anthony! I just received GCP credits from the competition and just finished setting the (preemptible) GPU. So I can also test more fancy methods next week.",
      "votes": null
    },
    {
      "id": "591326",
      "postDate": "08/03/2019 13:59:11",
      "content": "<p>OK, now also added the code for doing inference using Google AutoML:\n<a href=\"https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl\">https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl</a></p>\n\n<p>I didn't do my strongest model to avoid messing with the leaderboard (I picked a model that performs a little below to strongest performing kernel). To change models, all one needs to do is change the <strong>model_id</strong>.   </p>",
      "rawMarkdown": "OK, now also added the code for doing inference using Google AutoML:\n[https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl](https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl)\n\nI didn't do my strongest model to avoid messing with the leaderboard (I picked a model that performs a little below to strongest performing kernel). To change models, all one needs to do is change the **model_id**.",
      "votes": null
    },
    {
      "id": "600477",
      "postDate": "08/16/2019 06:32:19",
      "content": "<p>Managed to jump to 0.435 by exploiting the structure of the structure of the data (mentioned <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-596369\">here</a>). </p>",
      "rawMarkdown": "Managed to jump to 0.435 by exploiting the structure of the structure of the data (mentioned [here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-596369)).",
      "votes": null
    },
    {
      "id": "602122",
      "postDate": "08/18/2019 16:00:58",
      "content": "<p>Thanks for this <a href=\"/antgoldbloom\">@antgoldbloom</a> ! Good to know that it takes &gt;24h to train a performant model. I estimate that It'll take me at least 12hours to train an efficientnet-3 on 512px RGB images (two sites) that will score &gt;0.3 LB. I am not using mixed precision however and I use 1*Tesla V100.</p>\n\n<p>I wonder how much time it could take to train a gold-medal model &amp; with what resources and techniques - any ideas on this?</p>\n\n<p>This can get very 💰 (money) consuming if you don't have access to free computing resources.</p>\n\n<p>Update: It took me around 19hours to train a model that reached .315 accuracy with the architecture above.</p>",
      "rawMarkdown": "Thanks for this @antgoldbloom ! Good to know that it takes &gt;24h to train a performant model. I estimate that It'll take me at least 12hours to train an efficientnet-3 on 512px RGB images (two sites) that will score &gt;0.3 LB. I am not using mixed precision however and I use 1*Tesla V100.\n\nI wonder how much time it could take to train a gold-medal model &amp; with what resources and techniques - any ideas on this?\n\nThis can get very 💰 (money) consuming if you don't have access to free computing resources.\n\nUpdate: It took me around 19hours to train a model that reached .315 accuracy with the architecture above.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 583156,
      "author_name": "antgoldbloom",
      "author_url": "",
      "post_date": "07/24/2019 05:34:52",
      "content": "<h2>First Few Experiments</h2>\n\n<p>I started with the images created by <a href=\"/xhlulu\">@xhlulu</a>’s <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">kernel</a> (thank you <a href=\"/xhlulu\">@xhlulu</a>!). In that kernel, xhlulu@ converts the data to 224x224 RGB jpegs. </p>\n\n<p>By default, Google AutoML randomly splits between TRAIN, VALIDATION and TEST. And the predicted outputs labels and confidence scores for each image (I define a threshold and it outputs all labels and confidences above that confidence threshold). </p>\n\n<p>I made predictions for both sites and then picked the label that had a higher confidence score across both sites. Below is a table of my scores. (Note: AUC is the accuracy metric that Google AutoML Vision reports. Haven’t put the time into figuring out what it means but I assume it’s probably treating each class as binary on/off and then aggregating up.)</p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Public Leaderboard Score\n   \n  \n  \n   24\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.23\n   \n   0.101\n   \n  \n  \n   48\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.42\n   \n   \n   \n  \n  \n   72\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.5\n   \n   \n   \n  \n  \n   96\n   \n   224\n   \n   73030\n   \n   jpeg\n   \n   0.56\n   \n   0.199\n   \n  \n\n\n<p>I then made small modifications xhlulu@’s code to create 512x512 RGB pngs to test the impact of higher resolution images.</p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Leaderboard Score\n   \n  \n  \n   24\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.32\n   \n   \n   \n  \n  \n   48\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.31\n   \n   \n   \n  \n  \n   72\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.54\n   \n   0.254\n   \n  \n  \n   96\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.64\n   \n   0.272\n   \n  \n  \n   120\n   \n   512\n   \n   73030\n   \n   png\n   \n   0.68\n   \n   0.292\n   \n  \n\n\n<p>0.302 is my best score. It’s a result of a poorly thought out experiment, so I’m surprised it’s my best score. I trained a model for 144 hours that included the control images. But after inspecting the test-set predictions I realized my model was often picking labels &gt; 1108 (which are control labels that don’t appear in the test set). I ended up filtering the test-set predictions to go to the next most confident label when a label &gt;1108 was predicted. </p>\n\n\n  \n   Training Hours\n   \n   Resolution\n   \n   Number of Images\n   \n   Extension\n   \n   AutoML AUC\n   \n   Leaderboard Score\n   \n  \n  \n   24\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.34\n   \n   0.175\n   \n  \n  \n   48\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.54\n   \n   \n   \n  \n  \n   72\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.63\n   \n   \n   \n  \n  \n   96\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.68\n   \n   0.283\n   \n  \n  \n   120\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.71\n   \n   0.294\n   \n  \n  \n   144\n   \n   512\n   \n   yes\n   \n   png\n   \n   0.72\n   \n   0.302\n   \n  \n\n\n<p>Google AutoML Vision allows me to define my own TRAIN, VALIDATION and TEST splits. For my next experiments I’m going to start using this feature. I’m going to define my own splits to:\n1. Include the control data in TRAIN but not VALIDATION or TEST (to allow the model to use the control data for training but to discourage the model from predicting labels &gt;1108 on the competition test set)\n2. Split by experiment rather than randomly </p>",
      "votes": null,
      "replies": [
        {
          "id": 583285,
          "author_name": "lkhphuc",
          "author_url": "",
          "post_date": "07/24/2019 09:28:11",
          "content": "<p>Thank for sharing. Looking at your training hours, I think I should have more patience with my model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583489,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/24/2019 14:37:39",
          "content": "<p><a href=\"/lkhphuc\">@lkhphuc</a>, unlikely you need anywhere near as many training hours if you're not using AutoML. The upside of AutoML is that you don't have to set any hyperparameters and still get a decent model. One of the drawbacks is it's so compute intensive. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583525,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/24/2019 15:22:18",
          "content": "<p>Those results are pretty promising, considering that there was no architecture tuning whatsoever. The massive improvement in accuracy when you change the input from 224px jpegs to 512px pngs also seems to indicate that the compression format and size of the images plays an important factor in the accuracy.</p>\n\n<p>This is surprising in some way, since most of the classical architectures are designed to receive as input image sizes between 200 and 300, and greater image sizes usually requires training a filter that downsample the image into a size accepted by the network, or else it could result in OOM issues.</p>\n\n<p>If you still have the 512px png dataset laying around, could you possibly share it as a public dataset? That would save some of us spending hours generating the dataset locally ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583537,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/24/2019 15:33:10",
          "content": "<p><a href=\"/xhlulu\">@xhlulu</a>, I actually regret converting to png. It would have been a purer comparison if I'd also used jpeg. </p>\n\n<p>I tried running the 512px conversion in a kernel but I didn't have quite enough compute time. Let me find out how I can share the 512px dataset with everyone. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583539,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/24/2019 15:34:30",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583759,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/25/2019 00:10:31",
          "content": "<p><a href=\"/xhlulu\">@xhlulu</a>, unfortunately can't share the dataset through the datasets platform. That'd allow users to access the dataset without having to accept the competitions rules. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 583158,
      "author_name": "antgoldbloom",
      "author_url": "",
      "post_date": "07/24/2019 05:40:40",
      "content": "<h2>Experience with Google AutoML Vision</h2>\n\n<p>I’m getting pretty decent results considering how naive my models are. I’ve found Google AutoML easy in many ways. To get a basic model all I need to do is:\n1.  upload the RGB images to Google Cloud Storage\n2.  upload a CSV file with a pointer to the GCS bucket and the target label</p>\n\n<p>There are a bunch of limitations with the products. My biggest issues so far have been:\n1. inability to add other metadata (e.g. would like to be able to add metadata on controls, plate id and position on plate). I can probably get around this using a post-processing step but it’d be nice to be able to add this metadata into the single model.\n2. there isn’t a feature that allows batch prediction on a large number of images. I work around this by hitting the prediction API 19897 times to generate my submission file. \n3. the model evaluation page is not very helpful for debugging model performance. <br>\n4. I had to use a lot of hours of training to get good results (you can see that the models keep improving with additional training time). This means I have to wait a long time and spend a lot of money.\n5. can’t change the loss function. Based on the forums, it seems like the loss function might be an important setting for this competition. </p>\n\n<p>There are some more minor frustrations that I had to workaround, which I’m happy to share if others are planning to try it out and want to learn from some of the frictions I encountered.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 583248,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/24/2019 08:09:33",
      "content": "<p>Hi Anthony, Nice to meet you!!! You are an inspiration ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 589262,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/31/2019 17:14:58",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> thanks! Looks like you're a few places ahead of me on the leaderboard. Watch out, I have a few new ideas to try 😏 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589452,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/31/2019 23:04:13",
          "content": "<p><a href=\"/antgoldbloom\">@antgoldbloom</a> That’s interesting Anthony! I just received GCP credits from the competition and just finished setting the (preemptible) GPU. So I can also test more fancy methods next week.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 583711,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "07/24/2019 21:29:38",
      "content": "<p>Very interesting to see your code. Can Google AutoML be run on kernels (of course, there is a running time limitation of 9h, but still)? Maybe you can publish on github? </p>\n\n<p>Does it give architecture and hyper parameters recommendations that can be used separately afterwards?</p>",
      "votes": null,
      "replies": [
        {
          "id": 589256,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/31/2019 17:12:11",
          "content": "<p><a href=\"/zaharch\">@zaharch</a>, sorry for the delay. I wanted to see if AutoML could be run out of Kernels. Turns out it can:\n<a href=\"https://www.kaggle.com/antgoldbloom/training-using-google-automl/\">https://www.kaggle.com/antgoldbloom/training-using-google-automl/</a></p>\n\n<p>The 9 hour limit for kernels is not an issue because my kernel kicks off a 24 hour training job and then stops. </p>\n\n<p>My next step is to share my inference code which sends the test images to the trained model and generates predicted labels. </p>\n\n<p>AutoML doesn't give architecture or hyper parameter recommendations. It just creates an end point that you can send images to and get back predicted labels (and the model's confidence in that label).  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591326,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "08/03/2019 13:59:11",
          "content": "<p>OK, now also added the code for doing inference using Google AutoML:\n<a href=\"https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl\">https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl</a></p>\n\n<p>I didn't do my strongest model to avoid messing with the leaderboard (I picked a model that performs a little below to strongest performing kernel). To change models, all one needs to do is change the <strong>model_id</strong>.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602122,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/18/2019 16:00:58",
          "content": "<p>Thanks for this <a href=\"/antgoldbloom\">@antgoldbloom</a> ! Good to know that it takes &gt;24h to train a performant model. I estimate that It'll take me at least 12hours to train an efficientnet-3 on 512px RGB images (two sites) that will score &gt;0.3 LB. I am not using mixed precision however and I use 1*Tesla V100.</p>\n\n<p>I wonder how much time it could take to train a gold-medal model &amp; with what resources and techniques - any ideas on this?</p>\n\n<p>This can get very 💰 (money) consuming if you don't have access to free computing resources.</p>\n\n<p>Update: It took me around 19hours to train a model that reached .315 accuracy with the architecture above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 585543,
      "author_name": "ttylacm",
      "author_url": "",
      "post_date": "07/27/2019 15:55:45",
      "content": "<p>Thank you for sharing this awesome tool! May I know what's your GPU resource?</p>",
      "votes": null,
      "replies": [
        {
          "id": 589258,
          "author_name": "antgoldbloom",
          "author_url": "",
          "post_date": "07/31/2019 17:13:40",
          "content": "<p><a href=\"/ttylacm\">@ttylacm</a>, don't know what hardware Google AutoML is using under the hood. It's invisible to the user. Good chance it's using TPUs though.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 600477,
      "author_name": "antgoldbloom",
      "author_url": "",
      "post_date": "08/16/2019 06:32:19",
      "content": "<p>Managed to jump to 0.435 by exploiting the structure of the structure of the data (mentioned <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-596369\">here</a>). </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "583151": "I’m a long-time competition lurker who's finally mustered the courage to join a competition 😅. I picked this competition because I find the topic interesting (if I was going to college today I’d probably pick biology). And also as a chance to try out [Google AutoML Vision](https://cloud.google.com/vision/automl/docs/). \n\nKaggle is part of Google now, which gives me exposure to some of the technologies that are generating excitement inside Google. AutoML is a tool with a lot of buzz. \n\nI hadn’t paid much attention to the Automated Machine Learning Tools (AMLTs) until the KaggleDays competition in San Francisco this past April. Two AMLTs, Google AutoML and H20, ended up participating in the hackathon and getting top 10 performances (here’s the [Google AI writeup](https://ai.googleblog.com/2019/05/an-end-to-end-automl-solution-for.html) of their performance). While I believe it requires human creativity to do problem setup and domain-specific feature engineering, KaggleDays left me wondering whether humans need to be doing generic feature engineering, architecture selection, picking activation functions and setting learning rates. \n\nI plan to share details of my experiments and experiences with Google AutoML Vision in this thread. I’m also open to suggestions from others in the community for things I should try. But bear with me if I take some-time to respond: like many other Kagglers I have a busy day job.",
    "583156": "## First Few Experiments\n\nI started with the images created by @xhlulu’s [kernel](https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images) (thank you @xhlulu!). In that kernel, xhlulu@ converts the data to 224x224 RGB jpegs. \n\nBy default, Google AutoML randomly splits between TRAIN, VALIDATION and TEST. And the predicted outputs labels and confidence scores for each image (I define a threshold and it outputs all labels and confidences above that confidence threshold). \n\nI made predictions for both sites and then picked the label that had a higher confidence score across both sites. Below is a table of my scores. (Note: AUC is the accuracy metric that Google AutoML Vision reports. Haven’t put the time into figuring out what it means but I assume it’s probably treating each class as binary on/off and then aggregating up.)\n\n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Public Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.23\n   </td>\n   <td>0.101\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.42\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.5\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>224\n   </td>\n   <td>73030\n   </td>\n   <td>jpeg\n   </td>\n   <td>0.56\n   </td>\n   <td>0.199\n   </td>\n  </tr>\n</tbody></table>\n\n\nI then made small modifications xhlulu@’s code to create 512x512 RGB pngs to test the impact of higher resolution images.\n\n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.32\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.31\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.54\n   </td>\n   <td>0.254\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.64\n   </td>\n   <td>0.272\n   </td>\n  </tr>\n  <tr>\n   <td>120\n   </td>\n   <td>512\n   </td>\n   <td>73030\n   </td>\n   <td>png\n   </td>\n   <td>0.68\n   </td>\n   <td>0.292\n   </td>\n  </tr>\n</tbody></table>\n\n0.302 is my best score. It’s a result of a poorly thought out experiment, so I’m surprised it’s my best score. I trained a model for 144 hours that included the control images. But after inspecting the test-set predictions I realized my model was often picking labels &gt; 1108 (which are control labels that don’t appear in the test set). I ended up filtering the test-set predictions to go to the next most confident label when a label &gt;1108 was predicted. \n\n<table>\n  <tbody><tr>\n   <td>Training Hours\n   </td>\n   <td>Resolution\n   </td>\n   <td>Number of Images\n   </td>\n   <td>Extension\n   </td>\n   <td>AutoML AUC\n   </td>\n   <td>Leaderboard Score\n   </td>\n  </tr>\n  <tr>\n   <td>24\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.34\n   </td>\n   <td>0.175\n   </td>\n  </tr>\n  <tr>\n   <td>48\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.54\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>72\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.63\n   </td>\n   <td>\n   </td>\n  </tr>\n  <tr>\n   <td>96\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.68\n   </td>\n   <td>0.283\n   </td>\n  </tr>\n  <tr>\n   <td>120\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.71\n   </td>\n   <td>0.294\n   </td>\n  </tr>\n  <tr>\n   <td>144\n   </td>\n   <td>512\n   </td>\n   <td>yes\n   </td>\n   <td>png\n   </td>\n   <td>0.72\n   </td>\n   <td>0.302\n   </td>\n  </tr>\n</tbody></table>\n\nGoogle AutoML Vision allows me to define my own TRAIN, VALIDATION and TEST splits. For my next experiments I’m going to start using this feature. I’m going to define my own splits to:\n1. Include the control data in TRAIN but not VALIDATION or TEST (to allow the model to use the control data for training but to discourage the model from predicting labels &gt;1108 on the competition test set)\n2. Split by experiment rather than randomly",
    "583158": "## Experience with Google AutoML Vision\n\nI’m getting pretty decent results considering how naive my models are. I’ve found Google AutoML easy in many ways. To get a basic model all I need to do is:\n1.  upload the RGB images to Google Cloud Storage\n2.  upload a CSV file with a pointer to the GCS bucket and the target label\n\nThere are a bunch of limitations with the products. My biggest issues so far have been:\n1. inability to add other metadata (e.g. would like to be able to add metadata on controls, plate id and position on plate). I can probably get around this using a post-processing step but it’d be nice to be able to add this metadata into the single model.\n2. there isn’t a feature that allows batch prediction on a large number of images. I work around this by hitting the prediction API 19897 times to generate my submission file. \n3. the model evaluation page is not very helpful for debugging model performance.  \n4. I had to use a lot of hours of training to get good results (you can see that the models keep improving with additional training time). This means I have to wait a long time and spend a lot of money.\n5. can’t change the loss function. Based on the forums, it seems like the loss function might be an important setting for this competition. \n\nThere are some more minor frustrations that I had to workaround, which I’m happy to share if others are planning to try it out and want to learn from some of the frictions I encountered.",
    "583248": "Hi Anthony, Nice to meet you!!! You are an inspiration ;)",
    "583285": "Thank for sharing. Looking at your training hours, I think I should have more patience with my model.",
    "583489": "lkhphuc, unlikely you need anywhere near as many training hours if you're not using AutoML. The upside of AutoML is that you don't have to set any hyperparameters and still get a decent model. One of the drawbacks is it's so compute intensive.",
    "583525": "Those results are pretty promising, considering that there was no architecture tuning whatsoever. The massive improvement in accuracy when you change the input from 224px jpegs to 512px pngs also seems to indicate that the compression format and size of the images plays an important factor in the accuracy.\n\nThis is surprising in some way, since most of the classical architectures are designed to receive as input image sizes between 200 and 300, and greater image sizes usually requires training a filter that downsample the image into a size accepted by the network, or else it could result in OOM issues.\n\nIf you still have the 512px png dataset laying around, could you possibly share it as a public dataset? That would save some of us spending hours generating the dataset locally ;)",
    "583537": "xhlulu, I actually regret converting to png. It would have been a purer comparison if I'd also used jpeg. \n\nI tried running the 512px conversion in a kernel but I didn't have quite enough compute time. Let me find out how I can share the 512px dataset with everyone.",
    "583539": "Thanks!",
    "583711": "Very interesting to see your code. Can Google AutoML be run on kernels (of course, there is a running time limitation of 9h, but still)? Maybe you can publish on github? \n\nDoes it give architecture and hyper parameters recommendations that can be used separately afterwards?",
    "583759": "xhlulu, unfortunately can't share the dataset through the datasets platform. That'd allow users to access the dataset without having to accept the competitions rules.",
    "585543": "Thank you for sharing this awesome tool! May I know what's your GPU resource?",
    "589256": "zaharch, sorry for the delay. I wanted to see if AutoML could be run out of Kernels. Turns out it can:\nhttps://www.kaggle.com/antgoldbloom/training-using-google-automl/\n\nThe 9 hour limit for kernels is not an issue because my kernel kicks off a 24 hour training job and then stops. \n\nMy next step is to share my inference code which sends the test images to the trained model and generates predicted labels. \n\nAutoML doesn't give architecture or hyper parameter recommendations. It just creates an end point that you can send images to and get back predicted labels (and the model's confidence in that label).",
    "589258": "ttylacm, don't know what hardware Google AutoML is using under the hood. It's invisible to the user. Good chance it's using TPUs though.",
    "589262": "ratthachat thanks! Looks like you're a few places ahead of me on the leaderboard. Watch out, I have a few new ideas to try 😏",
    "589452": "antgoldbloom That’s interesting Anthony! I just received GCP credits from the competition and just finished setting the (preemptible) GPU. So I can also test more fancy methods next week.",
    "591326": "OK, now also added the code for doing inference using Google AutoML:\n[https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl](https://www.kaggle.com/antgoldbloom/doing-inference-using-google-automl)\n\nI didn't do my strongest model to avoid messing with the leaderboard (I picked a model that performs a little below to strongest performing kernel). To change models, all one needs to do is change the **model_id**.",
    "600477": "Managed to jump to 0.435 by exploiting the structure of the structure of the data (mentioned [here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-596369)).",
    "602122": "Thanks for this @antgoldbloom ! Good to know that it takes &gt;24h to train a performant model. I estimate that It'll take me at least 12hours to train an efficientnet-3 on 512px RGB images (two sites) that will score &gt;0.3 LB. I am not using mixed precision however and I use 1*Tesla V100.\n\nI wonder how much time it could take to train a gold-medal model &amp; with what resources and techniques - any ideas on this?\n\nThis can get very 💰 (money) consuming if you don't have access to free computing resources.\n\nUpdate: It took me around 19hours to train a model that reached .315 accuracy with the architecture above."
  },
  "source": "meta"
}