TechForge

December 3, 2024

  • Meta showcased its approach to scaling large language models.
  • Optimised GPU clusters and custom algorithms shape how Meta trains next-generation models.

At QCon San Francisco 2024, Meta’s Ye (Charlotte) Qi took the stage to discuss the challenges of running LLMs at scale. Her presentation shed light on the complexities of managing massive models in real-world systems, emphasising the obstacles posed by their sheer size, intricate hardware requirements, and demanding production environments. As she put it, the AI boom resembles an “AI Gold Rush,” where organisations seek innovation while confronting unprecedented technical hurdles.

The transition to generative AI has fundamentally changed the scale and scope of Meta’s AI infrastructure. Historically, the company focused on training a variety of smaller models, such as recommendation systems powering feeds and rankings. These models required relatively fewer GPUs to process datasets and deliver predictions. However, generative AI has brought a paradigm shift, with fewer but significantly larger jobs demanding an new approach to infrastructure design.

The evolution has introduced distinct challenges. For example, as the number of GPUs required for a single job increases, so does the likelihood of hardware failure. To ensure that such events don’t negatively impact service, the company has made improvements throughout its hardware stack and software running on it. Hardware reliability has become a cornerstone of Meta’s strategy, with rigorous testing protocols and automated systems in place to detect and address failures quickly.

At the same time, fast recovery mechanisms have been developed to minimise downtime, including techniques to preserve and restore training states in the event of hardware outage.

Qi also delved into how Meta has had to rethink GPU connectivity to support LLM training at scale. Training such models necessitates the synchronised transfer of massive amounts of data across GPUs, necessitating a reliable and high-speed network infrastructure. A single bottleneck can spread throughout the system, significantly slowing down operations.

To address this, Meta uses customised algorithms and optimised data exchange protocols to assure seamless communication across thousands of GPUs.

The task of fitting massive models onto hardware is equally difficult. Models with billions of parameters often exceed the capacity of a single GPU, making techniques such as tensor parallelism and pipeline parallelism necessary. Qi stated that aligning model architectures with hardware constraints is vital, as even minor mismatches can cause considerable performance decline.

She advised against relying on generic tools for this purpose, urging teams to adopt specialised runtimes for inference and tailor their approach to the unique demands of each model.

Meta claims its innovations go beyond hardware and network infrastructure. The software layer is “critical” to overall efficiency and adaptability. Scheduling algorithms dynamically allocate resources to meet the ephemeral needs of training jobs, optimising for efficiency and cost-effectiveness.

See also:

As Qi explored the transition from prototype to production, she emphasised the unpredictable nature of real-world conditions. Production workloads are frequently volatile, requiring systems to balance latency, reliability, and cost. Meta has fine-tuned its data centres to maximise compute density while balancing constraints determined by fixed factors like power, cooling, and network infrastructure. Its initiatives have involved relocating non-essential services to free up resources and packing GPU racks to maximise compute capability in each data hall.

Network infrastructure has been a particularly complex area for Meta. When faced with the decision between RoCE and InfiniBand fabrics, the company chose to deploy two 24,000-GPU clusters, one for each fabric type. This method let Meta compare the performance and operational viability of each network solution before training its next-generation Llama 3 model on both clusters.

Despite fundamental differences between the two fabrics, Meta achieved near-equivalent performance by optimising communication patterns, collective algorithms, and load balancing.

Meta’s journey to scale LLMs also underscores the importance of storage solutions. The vast datasets required for training demand high-capacity, high-speed storage systems. Qi said her company’s innovations in this area have enabled Meta to streamline data retrieval and checkpointing processes, guaranteeing that training jobs recover quickly from interruptions.

Throughout her presentation, Qi emphasised the importance of taking a holistic approach to scaling AI systems. While technical expertise is important, she implied that success also requires taking a step back for greater objectivity. By focusing resources on efforts that deliver long-term value, organisations can refine their systems to meet the demands of an evolving AI landscape.

Scaling LLMs is not just about technological breakthroughs, she said, but also about strategy, collaboration, and a focus on real-world impact.

Want to learn more about cybersecurity and the cloud from industry leaders? Check out Cyber Security & Cloud Expo taking place in Amsterdam, California, and London. Explore other upcoming enterprise technology events and webinars powered by TechForge here.

Author

  • As a tech journalist, Zul focuses on topics including cloud computing, cybersecurity, and disruptive technology in the enterprise industry. He has expertise in moderating webinars and presenting content on video, in addition to having a background in networking technology.

    View all posts

About the Author

Muhammad Zulhusni

As a tech journalist, Zul focuses on topics including cloud computing, cybersecurity, and disruptive technology in the enterprise industry. He has expertise in moderating webinars and presenting content on video, in addition to having a background in networking technology.

Related

August 24, 2026

August 11, 2026

August 10, 2026

August 5, 2026

Join our Community

Subscribe now to get all our premium content and latest tech news delivered straight to your inbox

Popular

12371 view(s)
11427 view(s)
7693 view(s)
5372 view(s)

Subscribe

All our premium content and latest tech news delivered straight to your inbox

This field is for validation purposes and should be left unchanged.
Name(Required)