{
  "id": 435577,
  "title": "Tensor Tile, Tensor Shard, Simulated Annealing and Langevin Dynamics. ",
  "url": "/competitions/predict-ai-model-runtime/discussion/435577",
  "author_name": "",
  "post_date": "2023-08-30T04:00:35.637736200Z",
  "votes": 51,
  "comment_count": 15,
  "views": 0,
  "content": "<h1>Tensor Tiling and Tensor Sharding</h1>\n<p>\"Both techniques used in parallel training of machine learning models, but they serve different purposes.\"</p>\n<p>\"Tensor tiling is a technique used to optimize the performance of tensor operations by partitioning the tensor into smaller, fixed-size tiles that can be loaded into memory and processed more efficiently.\"</p>\n<p>\"Tensor sharding, on the other hand, is a technique used to distribute the computation of large tensors across multiple devices or machines in a distributed system. The tensor is divided into smaller pieces, or shards, and each shard is processed independently on different devices.\"</p>\n<p>\"Both techniques can be used in conjunction with XLA (Accelerated Linear Algebra) and Hlo (High-Level Optimizer) technologies to optimize the computation graph used in deep learning training. GSPMD (gated synchronous parallelism data parallelism) is a specific parallel training approach that leverages these technologies and techniques to efficiently distribute the data and computations required for training across multiple devices or machines.\"</p>\n<p>Answered by Joe Apr 12 at 10:38</p>\n<p><a href=\"https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\" target=\"_blank\">https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation</a></p>\n<h1>Simulated Annealing</h1>\n<p>\"Simulated annealing is a method for solving unconstrained and bound-constrained optimization problems. The method models the physical process of heating a material and then slowly lowering the temperature to decrease defects, thus minimizing the system energy.\"</p>\n<p>\"At each iteration of the simulated annealing algorithm, a new point is randomly generated. The distance of the new point from the current point, or the extent of the search, is based on a probability distribution with a scale proportional to the temperature.\"</p>\n<p>\"The algorithm accepts all new points that lower the objective, but also, with a certain probability, points that raise the objective. By accepting points that raise the objective, the algorithm avoids being trapped in local minima, and is able to explore globally for more possible solutions. An annealing schedule is selected to systematically decrease the temperature as the algorithm proceeds. As the temperature decreases, the algorithm reduces the extent of its search to converge to a minimum.\"</p>\n<p><a href=\"https://www.mathworks.com/help/gads/what-is-simulated-annealing.html\" target=\"_blank\">https://www.mathworks.com/help/gads/what-is-simulated-annealing.html</a></p>\n<h1>Stochastic Gradient Langevin Dynamics</h1>\n<p>Bayesian Learning via Stochastic Gradient Langevin Dynamics</p>\n<p>Authors: Max Welling and Yee Whye Teh</p>\n<p>\"In that paper the authors proposed a new framework for learning from large scale datasets based on iterative learning from small mini-batches. By adding the right amount of noise to a standard stochastic gradient optimization algorithm we show that the iterates will converge to samples from the true posterior distribution as we anneal the stepsize.\"</p>\n<p>\"This seamless transition between optimization and Bayesian posterior sampling provides an inbuilt protection against overfitting. The authors also proposed a practical method for Monte Carlo estimates of posterior statistics which monitors a “sampling threshold” and collects samples after it has been surpassed. They applied the method to three models: a mixture of Gaussians, logistic regression and ICA with natural gradients.\"</p>\n<p>\"Given the similarities between stochastic gradient algorithms and Langevin dynamics, it is natural to consider combining ideas from the two approaches. This allows efficient use of large datasets while allowing for parameter uncertainty to be captured in a Bayesian manner.</p>\n<p>\"The approach is straightforward: use Robbins-Monro stochastic gradients, add an amount of Gaussian noise balanced with the step size used, and allow step sizes to go to zero.\"</p>\n<p><a href=\"https://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf\" target=\"_blank\">https://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf</a></p>",
  "messages": [
    {
      "id": "2414992",
      "postDate": "08/30/2023 04:00:35",
      "content": "<h1>Tensor Tiling and Tensor Sharding</h1>\n<p>\"Both techniques used in parallel training of machine learning models, but they serve different purposes.\"</p>\n<p>\"Tensor tiling is a technique used to optimize the performance of tensor operations by partitioning the tensor into smaller, fixed-size tiles that can be loaded into memory and processed more efficiently.\"</p>\n<p>\"Tensor sharding, on the other hand, is a technique used to distribute the computation of large tensors across multiple devices or machines in a distributed system. The tensor is divided into smaller pieces, or shards, and each shard is processed independently on different devices.\"</p>\n<p>\"Both techniques can be used in conjunction with XLA (Accelerated Linear Algebra) and Hlo (High-Level Optimizer) technologies to optimize the computation graph used in deep learning training. GSPMD (gated synchronous parallelism data parallelism) is a specific parallel training approach that leverages these technologies and techniques to efficiently distribute the data and computations required for training across multiple devices or machines.\"</p>\n<p>Answered by Joe Apr 12 at 10:38</p>\n<p><a href=\"https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\" target=\"_blank\">https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation</a></p>\n<h1>Simulated Annealing</h1>\n<p>\"Simulated annealing is a method for solving unconstrained and bound-constrained optimization problems. The method models the physical process of heating a material and then slowly lowering the temperature to decrease defects, thus minimizing the system energy.\"</p>\n<p>\"At each iteration of the simulated annealing algorithm, a new point is randomly generated. The distance of the new point from the current point, or the extent of the search, is based on a probability distribution with a scale proportional to the temperature.\"</p>\n<p>\"The algorithm accepts all new points that lower the objective, but also, with a certain probability, points that raise the objective. By accepting points that raise the objective, the algorithm avoids being trapped in local minima, and is able to explore globally for more possible solutions. An annealing schedule is selected to systematically decrease the temperature as the algorithm proceeds. As the temperature decreases, the algorithm reduces the extent of its search to converge to a minimum.\"</p>\n<p><a href=\"https://www.mathworks.com/help/gads/what-is-simulated-annealing.html\" target=\"_blank\">https://www.mathworks.com/help/gads/what-is-simulated-annealing.html</a></p>\n<h1>Stochastic Gradient Langevin Dynamics</h1>\n<p>Bayesian Learning via Stochastic Gradient Langevin Dynamics</p>\n<p>Authors: Max Welling and Yee Whye Teh</p>\n<p>\"In that paper the authors proposed a new framework for learning from large scale datasets based on iterative learning from small mini-batches. By adding the right amount of noise to a standard stochastic gradient optimization algorithm we show that the iterates will converge to samples from the true posterior distribution as we anneal the stepsize.\"</p>\n<p>\"This seamless transition between optimization and Bayesian posterior sampling provides an inbuilt protection against overfitting. The authors also proposed a practical method for Monte Carlo estimates of posterior statistics which monitors a “sampling threshold” and collects samples after it has been surpassed. They applied the method to three models: a mixture of Gaussians, logistic regression and ICA with natural gradients.\"</p>\n<p>\"Given the similarities between stochastic gradient algorithms and Langevin dynamics, it is natural to consider combining ideas from the two approaches. This allows efficient use of large datasets while allowing for parameter uncertainty to be captured in a Bayesian manner.</p>\n<p>\"The approach is straightforward: use Robbins-Monro stochastic gradients, add an amount of Gaussian noise balanced with the step size used, and allow step sizes to go to zero.\"</p>\n<p><a href=\"https://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf\" target=\"_blank\">https://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf</a></p>",
      "rawMarkdown": "#Tensor Tiling and Tensor Sharding\n\n\"Both techniques used in parallel training of machine learning models, but they serve different purposes.\"\n\n\"Tensor tiling is a technique used to optimize the performance of tensor operations by partitioning the tensor into smaller, fixed-size tiles that can be loaded into memory and processed more efficiently.\"\n\n\"Tensor sharding, on the other hand, is a technique used to distribute the computation of large tensors across multiple devices or machines in a distributed system. The tensor is divided into smaller pieces, or shards, and each shard is processed independently on different devices.\"\n\n\"Both techniques can be used in conjunction with XLA (Accelerated Linear Algebra) and Hlo (High-Level Optimizer) technologies to optimize the computation graph used in deep learning training. GSPMD (gated synchronous parallelism data parallelism) is a specific parallel training approach that leverages these technologies and techniques to efficiently distribute the data and computations required for training across multiple devices or machines.\"\n\nAnswered by Joe Apr 12 at 10:38\n\nhttps://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\n\n#Simulated Annealing\n\n\"Simulated annealing is a method for solving unconstrained and bound-constrained optimization problems. The method models the physical process of heating a material and then slowly lowering the temperature to decrease defects, thus minimizing the system energy.\"\n\n\"At each iteration of the simulated annealing algorithm, a new point is randomly generated. The distance of the new point from the current point, or the extent of the search, is based on a probability distribution with a scale proportional to the temperature.\"\n\n\"The algorithm accepts all new points that lower the objective, but also, with a certain probability, points that raise the objective. By accepting points that raise the objective, the algorithm avoids being trapped in local minima, and is able to explore globally for more possible solutions. An annealing schedule is selected to systematically decrease the temperature as the algorithm proceeds. As the temperature decreases, the algorithm reduces the extent of its search to converge to a minimum.\"\n\nhttps://www.mathworks.com/help/gads/what-is-simulated-annealing.html\n\n#Stochastic Gradient Langevin Dynamics\n\nBayesian Learning via Stochastic Gradient Langevin Dynamics\n\nAuthors: Max Welling and Yee Whye Teh\n\n\"In that paper the authors proposed a new framework for learning from large scale datasets based on iterative learning from small mini-batches. By adding the right amount of noise to a standard stochastic gradient optimization algorithm we show that the iterates will converge to samples from the true posterior distribution as we anneal the stepsize.\"\n\n\"This seamless transition between optimization and Bayesian posterior sampling provides an inbuilt protection against overfitting. The authors also proposed a practical method for Monte Carlo estimates of posterior statistics which monitors a “sampling threshold” and collects samples after it has been surpassed. They applied the method to three models: a mixture of Gaussians, logistic regression and ICA with natural gradients.\"\n\n\"Given the similarities between stochastic gradient algorithms and Langevin dynamics, it is natural to consider combining ideas from the two approaches. This allows efficient use of large datasets while allowing for parameter uncertainty to be captured in a Bayesian manner.\n\n\"The approach is straightforward: use Robbins-Monro stochastic gradients, add an amount of Gaussian noise balanced with the step size used, and allow step sizes to go to zero.\"\n\nhttps://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf",
      "votes": null
    },
    {
      "id": "2415185",
      "postDate": "08/30/2023 06:57:06",
      "content": "<p>Great content to start this competition.Thank You !</p>",
      "rawMarkdown": "Great content to start this competition.Thank You !",
      "votes": null
    },
    {
      "id": "2415576",
      "postDate": "08/30/2023 13:07:51",
      "content": "<p>We're going to learn a lot on this competition since everything is different with npz files, Tensor tiling, XLA and much more.  </p>",
      "rawMarkdown": "We're going to learn a lot on this competition since everything is different with npz files, Tensor tiling, XLA and much more.",
      "votes": null
    },
    {
      "id": "2415900",
      "postDate": "08/30/2023 17:23:24",
      "content": "<p>Thank you for the good educational summary about these topics!</p>",
      "rawMarkdown": "Thank you for the good educational summary about these topics!",
      "votes": null
    },
    {
      "id": "2415915",
      "postDate": "08/30/2023 17:30:45",
      "content": "<p>Since I'm a beginner, it was the first time that I read about such subjects Samihaija.</p>",
      "rawMarkdown": "Since I'm a beginner, it was the first time that I read about such subjects Samihaija.",
      "votes": null
    },
    {
      "id": "2416504",
      "postDate": "08/31/2023 05:15:47",
      "content": "<p><a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> Thank you for the details analysis.</p>",
      "rawMarkdown": "mpwolke Thank you for the details analysis.",
      "votes": null
    },
    {
      "id": "2420074",
      "postDate": "09/02/2023 11:54:59",
      "content": "<p>Thanks for sharing the information on Tensor Tiling and Tensor Sharding.<br>\nWe can also observe the multi-core computing with many CPUs.</p>",
      "rawMarkdown": "Thanks for sharing the information on Tensor Tiling and Tensor Sharding.\nWe can also observe the multi-core computing with many CPUs.",
      "votes": null
    },
    {
      "id": "2420310",
      "postDate": "09/02/2023 14:43:31",
      "content": "<p>I've only learned those concepts now CR S Kumar.<br>\nBy the way, you're doing great on this competition. Good luck!</p>",
      "rawMarkdown": "I've only learned those concepts now CR S Kumar.\nBy the way, you're doing great on this competition. Good luck!",
      "votes": null
    },
    {
      "id": "2420331",
      "postDate": "09/02/2023 14:51:57",
      "content": "<p>Thank Miah. I'm glad to hear your words.</p>",
      "rawMarkdown": "Thank Miah. I'm glad to hear your words.",
      "votes": null
    },
    {
      "id": "2420547",
      "postDate": "09/02/2023 17:23:04",
      "content": "<p>Do tensor tiles or tensor shards play a role in improving the efficiency or performance of deep learning models?</p>",
      "rawMarkdown": "Do tensor tiles or tensor shards play a role in improving the efficiency or performance of deep learning models?",
      "votes": null
    },
    {
      "id": "2420559",
      "postDate": "09/02/2023 17:35:14",
      "content": "<p>Hi Miah,</p>\n<p>As I wrote (copied from StackOverflow on the 1st §):</p>\n<p>\"Both techniques can be used in conjunction with XLA and HLO technologies to optimize the computation graph used in deep learning training.\"</p>\n<p><a href=\"https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\" target=\"_blank\">https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation</a></p>",
      "rawMarkdown": "Hi Miah,\n\nAs I wrote (copied from StackOverflow on the 1st §):\n\n\"Both techniques can be used in conjunction with XLA and HLO technologies to optimize the computation graph used in deep learning training.\"\n\nhttps://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation",
      "votes": null
    },
    {
      "id": "2420577",
      "postDate": "09/02/2023 17:51:08",
      "content": "<p>Thanks for replying, I heard those terms first time in your article and tried to learn. Keep writing more articles I learned a lot from you.</p>",
      "rawMarkdown": "Thanks for replying, I heard those terms first time in your article and tried to learn. Keep writing more articles I learned a lot from you.",
      "votes": null
    },
    {
      "id": "2423238",
      "postDate": "09/04/2023 14:23:05",
      "content": "<p>Thanks for sharing this information <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> . I am currently participating in this competition and this information will definitely help.</p>",
      "rawMarkdown": "Thanks for sharing this information @mpwolke . I am currently participating in this competition and this information will definitely help.",
      "votes": null
    },
    {
      "id": "2423286",
      "postDate": "09/04/2023 14:43:39",
      "content": "<p>You're doing great on this competition. I'm glad that you found it useful.</p>",
      "rawMarkdown": "You're doing great on this competition. I'm glad that you found it useful.",
      "votes": null
    },
    {
      "id": "2425162",
      "postDate": "09/05/2023 17:30:51",
      "content": "<p>Very clean content. Thank you for the explanations <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> !</p>",
      "rawMarkdown": "Very clean content. Thank you for the explanations @mpwolke !",
      "votes": null
    },
    {
      "id": "2425180",
      "postDate": "09/05/2023 17:36:25",
      "content": "<p>Merci beaucoup Thibaud. <br>\nI'm glad that you found it clean.</p>",
      "rawMarkdown": "Merci beaucoup Thibaud. \nI'm glad that you found it clean.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2415185,
      "author_name": "mysticshadow",
      "author_url": "",
      "post_date": "08/30/2023 06:57:06",
      "content": "<p>Great content to start this competition.Thank You !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2415576,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/30/2023 13:07:51",
          "content": "<p>We're going to learn a lot on this competition since everything is different with npz files, Tensor tiling, XLA and much more.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2415900,
      "author_name": "samihaija",
      "author_url": "",
      "post_date": "08/30/2023 17:23:24",
      "content": "<p>Thank you for the good educational summary about these topics!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2415915,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/30/2023 17:30:45",
          "content": "<p>Since I'm a beginner, it was the first time that I read about such subjects Samihaija.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2416504,
      "author_name": "faysalmiah1721758",
      "author_url": "",
      "post_date": "08/31/2023 05:15:47",
      "content": "<p><a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> Thank you for the details analysis.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2420331,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/02/2023 14:51:57",
          "content": "<p>Thank Miah. I'm glad to hear your words.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2420074,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "09/02/2023 11:54:59",
      "content": "<p>Thanks for sharing the information on Tensor Tiling and Tensor Sharding.<br>\nWe can also observe the multi-core computing with many CPUs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2420310,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/02/2023 14:43:31",
          "content": "<p>I've only learned those concepts now CR S Kumar.<br>\nBy the way, you're doing great on this competition. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2420547,
      "author_name": "faysalmiah1721758",
      "author_url": "",
      "post_date": "09/02/2023 17:23:04",
      "content": "<p>Do tensor tiles or tensor shards play a role in improving the efficiency or performance of deep learning models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2420559,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/02/2023 17:35:14",
          "content": "<p>Hi Miah,</p>\n<p>As I wrote (copied from StackOverflow on the 1st §):</p>\n<p>\"Both techniques can be used in conjunction with XLA and HLO technologies to optimize the computation graph used in deep learning training.\"</p>\n<p><a href=\"https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\" target=\"_blank\">https://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2420577,
              "author_name": "faysalmiah1721758",
              "author_url": "",
              "post_date": "09/02/2023 17:51:08",
              "content": "<p>Thanks for replying, I heard those terms first time in your article and tried to learn. Keep writing more articles I learned a lot from you.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2423238,
      "author_name": "rishabh15virgo",
      "author_url": "",
      "post_date": "09/04/2023 14:23:05",
      "content": "<p>Thanks for sharing this information <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> . I am currently participating in this competition and this information will definitely help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2423286,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/04/2023 14:43:39",
          "content": "<p>You're doing great on this competition. I'm glad that you found it useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2425162,
      "author_name": "theudbald",
      "author_url": "",
      "post_date": "09/05/2023 17:30:51",
      "content": "<p>Very clean content. Thank you for the explanations <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2425180,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/05/2023 17:36:25",
          "content": "<p>Merci beaucoup Thibaud. <br>\nI'm glad that you found it clean.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2414992": "#Tensor Tiling and Tensor Sharding\n\n\"Both techniques used in parallel training of machine learning models, but they serve different purposes.\"\n\n\"Tensor tiling is a technique used to optimize the performance of tensor operations by partitioning the tensor into smaller, fixed-size tiles that can be loaded into memory and processed more efficiently.\"\n\n\"Tensor sharding, on the other hand, is a technique used to distribute the computation of large tensors across multiple devices or machines in a distributed system. The tensor is divided into smaller pieces, or shards, and each shard is processed independently on different devices.\"\n\n\"Both techniques can be used in conjunction with XLA (Accelerated Linear Algebra) and Hlo (High-Level Optimizer) technologies to optimize the computation graph used in deep learning training. GSPMD (gated synchronous parallelism data parallelism) is a specific parallel training approach that leverages these technologies and techniques to efficiently distribute the data and computations required for training across multiple devices or machines.\"\n\nAnswered by Joe Apr 12 at 10:38\n\nhttps://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation\n\n#Simulated Annealing\n\n\"Simulated annealing is a method for solving unconstrained and bound-constrained optimization problems. The method models the physical process of heating a material and then slowly lowering the temperature to decrease defects, thus minimizing the system energy.\"\n\n\"At each iteration of the simulated annealing algorithm, a new point is randomly generated. The distance of the new point from the current point, or the extent of the search, is based on a probability distribution with a scale proportional to the temperature.\"\n\n\"The algorithm accepts all new points that lower the objective, but also, with a certain probability, points that raise the objective. By accepting points that raise the objective, the algorithm avoids being trapped in local minima, and is able to explore globally for more possible solutions. An annealing schedule is selected to systematically decrease the temperature as the algorithm proceeds. As the temperature decreases, the algorithm reduces the extent of its search to converge to a minimum.\"\n\nhttps://www.mathworks.com/help/gads/what-is-simulated-annealing.html\n\n#Stochastic Gradient Langevin Dynamics\n\nBayesian Learning via Stochastic Gradient Langevin Dynamics\n\nAuthors: Max Welling and Yee Whye Teh\n\n\"In that paper the authors proposed a new framework for learning from large scale datasets based on iterative learning from small mini-batches. By adding the right amount of noise to a standard stochastic gradient optimization algorithm we show that the iterates will converge to samples from the true posterior distribution as we anneal the stepsize.\"\n\n\"This seamless transition between optimization and Bayesian posterior sampling provides an inbuilt protection against overfitting. The authors also proposed a practical method for Monte Carlo estimates of posterior statistics which monitors a “sampling threshold” and collects samples after it has been surpassed. They applied the method to three models: a mixture of Gaussians, logistic regression and ICA with natural gradients.\"\n\n\"Given the similarities between stochastic gradient algorithms and Langevin dynamics, it is natural to consider combining ideas from the two approaches. This allows efficient use of large datasets while allowing for parameter uncertainty to be captured in a Bayesian manner.\n\n\"The approach is straightforward: use Robbins-Monro stochastic gradients, add an amount of Gaussian noise balanced with the step size used, and allow step sizes to go to zero.\"\n\nhttps://www.stats.ox.ac.uk/~teh/research/compstats/WelTeh2011a.pdf",
    "2415185": "Great content to start this competition.Thank You !",
    "2415576": "We're going to learn a lot on this competition since everything is different with npz files, Tensor tiling, XLA and much more.",
    "2415900": "Thank you for the good educational summary about these topics!",
    "2415915": "Since I'm a beginner, it was the first time that I read about such subjects Samihaija.",
    "2416504": "mpwolke Thank you for the details analysis.",
    "2420074": "Thanks for sharing the information on Tensor Tiling and Tensor Sharding.\nWe can also observe the multi-core computing with many CPUs.",
    "2420310": "I've only learned those concepts now CR S Kumar.\nBy the way, you're doing great on this competition. Good luck!",
    "2420331": "Thank Miah. I'm glad to hear your words.",
    "2420547": "Do tensor tiles or tensor shards play a role in improving the efficiency or performance of deep learning models?",
    "2420559": "Hi Miah,\n\nAs I wrote (copied from StackOverflow on the 1st §):\n\n\"Both techniques can be used in conjunction with XLA and HLO technologies to optimize the computation graph used in deep learning training.\"\n\nhttps://stackoverflow.com/questions/75994203/are-tensor-sharding-and-tensor-tilting-the-same-implementation",
    "2420577": "Thanks for replying, I heard those terms first time in your article and tried to learn. Keep writing more articles I learned a lot from you.",
    "2423238": "Thanks for sharing this information @mpwolke . I am currently participating in this competition and this information will definitely help.",
    "2423286": "You're doing great on this competition. I'm glad that you found it useful.",
    "2425162": "Very clean content. Thank you for the explanations @mpwolke !",
    "2425180": "Merci beaucoup Thibaud. \nI'm glad that you found it clean."
  },
  "source": "meta"
}